Data Engineer, PDS&T CMC
North Chicago, IL - USA
Job Summary
While the AI innovation race in Biopharma is focused on Drug discovery Product Development/ CMC represents the next barrier/ bottleneck. The complexity of biological systems the rigor of regulatory expectations the pace of pipeline growth and the enormous value at stake make this one of the highest-leverage domains for applied data science and AI in the entire pharmaceutical value chain.
We here at BTS - PDST are building a dedicated AI-native team that is driving cutting edge programs across early stage late stage and commercial product development to accelerate E2E product development and launch maximize yields of block buster products. Through our deep collaboration with PDST scientists we are boldly reimagining how AbbVie can bring our pipeline products and lifesaving drugs to patients faster safer and in cost effective manner fueled by AI.
This position is a highly technical AI-native role responsible for designing building and operating production-grade data pipelines and data products that power AI/ML analytics and automation across AbbVies CMC and manufacturing ecosystem.
This role is embedded inside PDST and works at the frontier of pharmaceutical data engineering. You will integrate and harmonize data from the full spectrum of manufacturing and development systems including MES historians LIMS QMS ERP and instrument platforms and transform it into reliable governed semantically rich data assets that data scientists process engineers and AI systems can actually use.
- Enterprise-scale scope: Enterprise-scale biologics portfolio spanning clinical commercial and lifecycle stages
- Building AI playbook for the future: First-in-AbbVie and first-in-biologics analytical approaches; you build the AI playbook for the future
- Growth and Impact: Direct impact on regulatory submissions commercial readiness and manufacturing decisions through deep cross-functional exposure to manufacturing quality regulatory and scientific leadership
- Mission: Every model you build helps ensure safe reliable medicines reach patients at scale
Responsibilities:
Data Ingestion & Integration
- Design and implement scalable robust data ingestion pipelines that connect CMC and manufacturing source systems including MES (Manufacturing Execution Systems) process historians LIMS QMS ERP platforms and instrument data sources to centralized and federated data environments.
- Build connectors adapters and integration layers that handle the heterogeneous data formats protocols and latency profiles characteristic of pharmaceutical manufacturing environments.
- Support both batch and real-time/streaming data patterns selecting appropriate architectures based on use case requirements.
Data Harmonization & Semantic Modeling
- Develop and maintain harmonized data models and ontologies that bring consistency to CMC and manufacturing data across sites systems and modalities.
- Execute semantic mapping efforts that align source system fields units and identifiers to enterprise data standards and scientific meaning.
- Collaborate with process scientists analytical chemists and manufacturing engineers to ensure data models accurately reflect domain reality.
Data Quality Observability & Governance
- Implement automated data quality controls validation frameworks and anomaly detection mechanisms across pipeline layers.
- Build and maintain data lineage documentation and metadata infrastructure enabling full traceability from source system to AI model input.
- Establish pipeline observability practices monitoring alerting SLA tracking to ensure data product reliability in production.
- Support data governance practices aligned with GxP requirements 21 CFR Part 11 and AbbVie data standards.
AI/ML Enablement & Data Product Development
- Architect and deliver governed versioned reusable data products purpose-built for AI/ML consumption including feature stores curated datasets and vector-ready data layers for RAG and LLM applications.
- Partner closely with data scientists ML engineers and process modelers to understand model data requirements and translate them into reliable scalable data infrastructure.
- Accelerate AI program delivery by eliminating data bottlenecks not by workarounds but by solving root causes structurally.
Platform & Operational Enablement
- Contribute to the design and evolution of PDSTs cloud-based data platform including lakehouse architecture data cataloging access control and compute infrastructure.
- Write and maintain infrastructure-as-code CI/CD pipelines and automated testing frameworks for data systems.
- Support platform onboarding of new CMC data domains and manufacturing sites ensuring consistent application of standards and patterns.
- Provide operational support for production data pipelines maintaining uptime and data freshness commitments.
Stakeholder Engagement & Scientific Leadership
- Influence technical decision-making without formal authority earning trust through scientific rigor transparent methodology and demonstrated business impact.
Qualifications :
Required:
- Bachelors Degree in Computer Science Data Engineering Information Systems Software Engineering Bioinformatics or a closely related technical field plus 2 years experience OR Masters Degree with 0 years experience.
- Respective years of hands-on experience designing and building enterprise-grade data pipelines integration workflows and data products in complex multi-source environments.
- Expert-level proficiency in Python for data engineering tasks pipeline development transformation logic data validation and automation.
- Strong SQL skills across modern analytical and transactional databases; comfort with both ANSI SQL and platform-specific dialects.
- Demonstrated experience with cloud data platforms (AWS Azure or GCP) and modern data stack components including tools such as dbt Spark Airflow Databricks Snowflake or equivalents.
- Develop ETL/ELT pipelines using tools such as Informatica Talend Apache NiFi and cloud-native services (e.g. AWS Glue Azure Data Factory).
- Implement master data management (MDM) metadata management and data cataloging solutions to ensure proper data lineage accessibility and compliance.
- Set and enforce standards for API development and data integration (REST GraphQL OData) enabling seamless integration using microservices architectures.
- Design logical physical and conceptual data models using modeling tools (e.g. Erwin PowerDesigner dbt).
- Ownership orientation: you define your own problem space drive solutions to completion and hold yourself accountable to outcomes not just outputs.
- Solution-architect instinct: you think before you build consider the full landscape of available approaches and choose tools based on fit-for-purpose reasoning rather than familiarity or trend.
- Scientific integrity: you build models you can explain defend and improve and you apply the same standard to the work of others.
- Influence through credibility: you earn the confidence of scientists engineers and quality professionals by being right being clear and being useful not by title or volume.
- Bias for impact: you are drawn to problems where the stakes are high and the analytical opportunity is real and you are energized rather than intimidated by ambiguity.
Preferred:
- Experience in pharmaceutical biotech or other regulated life sciences manufacturing environments.
- Familiarity with GxP data principles 21 CFR Part 11 compliance or data integrity requirements in regulated industries.
- Prior exposure to manufacturing source systems such as MES process historians (e.g. OSIsoft PI/AVEVA) LIMS QMS or ERP platforms
- Experience building data infrastructure for AI/ML programs including feature engineering pipelines model training datasets or vector/embedding data layers for RAG architectures.
Additional Information :
Applicable only to applicants applying to a position in any location with pay disclosure requirements under state or local law:
- The compensation range described below is the range of possible base pay compensation that the Company believes in good faith it will pay for this role at the time of this posting based on the job grade for this position. Individual compensation paid within this range will depend on many factors including geographic location and we may ultimately pay more or less than the posted range. This range may be modified in the future.
- We offer a comprehensive package of benefits including paid time off (vacation holidays sick) medical/dental/vision insurance and 401(k) to eligible employees.
- This job is eligible to participate in our short-term incentive programs.
Note: No amount of pay is considered to be wages or compensation until such amount is earned vested and determinable. The amount and availability of any bonus commission incentive benefits or any other form of compensation and benefits that are allocable to a particular employee remains in the Companys sole and absolute discretion unless and until paid and may be modified at the Companys sole and absolute discretion consistent with applicable law.
AbbVie is an equal opportunity employer and is committed to operating with integrity driving innovation transforming lives and serving our community. Equal Opportunity Employer/Veterans/Disabled.
US & Puerto Rico only - to learn more visit & Puerto Rico applicants seeking a reasonable accommodation click here to learn more:
No
Employment Type :
Full-time
About Company
AbbVie is a global biopharmaceutical company focused on creating medicines and solutions that put impact first for patients, communities, and our world. We aim to address complex health issues and enhance people's lives through our core therapeutic areas: immunology, oncology, neuro ... View more