Machine Learning Engineer, Infra, AI for Drug Discovery
New York City, NY - USA
Job Summary
A healthier future. Its what drives us to innovate. To continuously advance science and ensure everyone has access to the healthcare they need today and for generations to come. Creating a world where we all have more time with the people we love. Thats what makes us Roche.
Advances in AI data and computational sciences are transforming drug discovery and development. Roches Research and Early Development organisations at Genentech (gRED) and Pharma (pRED) have demonstrated how these technologies accelerate R&D leveraging data and novel computational models to drive impact. Seamless data sharing and access to models across gRED and pRED are essential to maximising these opportunities. The new Computational Sciences Center of Excellence (CoE) is a strategic unified group whose goal is to harness this transformative power of data and Artificial Intelligence (AI) to assist our scientists in both pRED and gRED to deliver more innovative and transformative medicines for patients worldwide.
The Opportunity
At Roches AI for Drug Discovery (AI4DD) group (Prescient Design) we are building the machine learning platforms that enable researchers and engineers to move models from experimentation into reliable scientific and production workflows. We are seeking a Machine Learning Infrastructure Engineer to help build and operate the platforms that support model deployment evaluation promotion monitoring and lifecycle management across the organization. This role will contribute to our model-serving platform and to the broader infrastructure required to make machine learning models easier to deploy scale observe and safely incorporate into scientific and agentic workflows.
The scope extends beyond LLM serving. You will work with a range of machine learning and scientific models including real-time and batch inference workloads GPU-backed services agentic applications and our in-silico drug discovery workflows. This is a hands-on engineering role for someone who enjoys writing and shipping production software across application code cloud infrastructure Kubernetes and distributed systems. Prior inference-platform experience is helpful but not required; prior experience in biotech or drug discovery is also helpful but not required; we value strong engineering fundamentals curiosity and the ability to take platform problems from design through production operation.
In this role you will:
Design implement ship and operate scalable model-serving infrastructure for machine learning scientific LLM and agentic workloads.
Help evolve our internal model deployment platform into a reliable self-service platform for teams across the organization.
Improve platform scalability and reliability including scale-to-zero faster model startup workload isolation traffic management and reduction of request failures and latency bottlenecks.
Build observability and operational tooling for model usage latency reliability resource consumption inference cost bottlenecks and service-level indicators.
Improve the usability of model deployment by developing validated configuration interfaces reusable deployment patterns APIs command-line tools and documentation.
Help converge real-time and batch inference workflows onto shared platform capabilities where appropriate.
Contribute to model lifecycle management infrastructure including model registration and versioning evaluation promotion and release gates monitoring environment progression and rollback.
Build event-driven integrations that connect model publication evaluation promotion deployment and retraining workflows.
Build consistent metrics and evaluation signals for understanding model cost quality reliability and fitness for downstream workflows.
Partner with machine learning data scientific and platform teams to translate requirements into maintainable solutions and remove infrastructure bottlenecks.
Own workstreams from design through implementation and production support using strong software-engineering practices including testing reviews documentation and incremental delivery.
Who You Are
BS or MS in Computer Science Engineering or a related technical field or equivalent practical experience.
3 years of relevant industry experience in software engineering infrastructure engineering platform engineering DevOps MLOps or a related area.
Strong Python programming skills and experience building and shipping maintainable production software services automation or developer tooling.
A demonstrated interest in hands-on implementation and production software delivery.
Experience designing deploying or operating cloud systems (preferably on AWS) using services such as EKS EC2 S3 IAM SQS SNS and CloudWatch.
Experience with containers Kubernetes Helm and IaC tools such as Terraform or Pulumi.
Experience with CI/CD Git-based development workflows automated testing and software release practices.
Ability to troubleshoot complex systems using metrics logs traces events and observability tools such as Datadog Prometheus Grafana or OpenTelemetry.
Understanding of distributed-systems concepts such as concurrency queuing retries timeouts idempotency backpressure and failure recovery.
Ability to gather requirements communicate technical tradeoffs and document systems for users and engineers with varied infrastructure experience.
Demonstrated ability to independently deliver practical incremental solutions while considering immediate needs and longer-term platform direction.
Preferred
Familiarity with model-serving or workflow-orchestration frameworks such as KServe Triton vLLM Ray Serve Prefect or Dagster.
Experience optimizing model startup time request throughput batching autoscaling or GPU utilization.
Familiarity with model registries experiment tracking model evaluation promotion workflows or MLOps platforms.
Experience building event-driven systems using queues event buses or workflow orchestrators.
Familiarity with online and offline model evaluation model-quality monitoring data drift or regression analysis.
Experience supporting scientific computing high-performance computing distributed training or large-scale data processing.
Strong interest in the life sciences and drug discovery.
Relocation benefits are NOT available for this job posting
The expected salary range for this position based on the primary location of California is $147600 - $274000 and for New York $141100 - 262100. Actual pay will be determined based on experience qualifications geographic location and other job-related factors permitted by law. A discretionary annual bonus may be available based on individual and Company performance. This position also qualifies for the benefits detailed at the link provided below.
#ComputationCoE
#tech4lifeComputationalScience
#tech4lifeAI
Genentech is an equal opportunity employer. It is our policy and practice to employ promote and otherwise treat any and all employees and applicants on the basis of merit qualifications and competence. The companys policy prohibits unlawful discrimination including but not limited to discrimination on the basis of Protected Veteran status individuals with disabilities status and consistent with all federal state or local laws.
If you have a disability and need an accommodation in relation to the online application process please contact us by completing this form Accommodations for Applicants.
Required Experience:
IC
About Company
As a pioneer in healthcare, we have been committed to improving lives since the company was founded in 1896 in Basel, Switzerland. Today, Roche creates innovative medicines and diagnostic tests that help millions of patients globally.