Enter a job title or keyword

Site Reliability Engineer ML, Apple Ads

Apple


Job Location:

New York City, NY - USA

Monthly Salary: Not provided by the employer
Posted: 1 October 2026 (Yesterday)
Application Deadline: 29 December 2026
Vacancies: 1 Vacancy

Job Summary

At Apple we focus deeply on our customers experience. Apple Ads brings this same approach to advertising helping people find exactly what theyre looking for and helping advertisers grow their technology powers ads and sponsorships across Apple Services including the App Store Apple News and MLS Season Pass. Everything we do is designed for trust connection and impact: We respect user privacy integrate advertising thoughtfully into the experience and deliver value for advertisers of all sizesfrom small app developers to big global brands. Because when advertising is done right it benefits Site Reliability Engineering team within Apple Ads ensures the reliability performance and availability of ML Platform and Services at scale. The team partners closely with Ads engineering data science and ML platform teams to enable product delivery through design configuration and automation of machine learning infrastructure powering Apple Ads are looking for a ML Platform Infrastructure Engineer to help build and evolve the next generation of Apple Ads machine learning platform enabling fast reliable and scalable operations across AWS-based environments supporting transactional and analytical workloads.

As a site reliability engineer in Apple Ads focused on machine learning you will own the health performance and scalability of large scale infrastructure powering ML training inference serving workloads and associated platform tooling. Your focus will be on building automation that eliminates manual processes improves platform resilience and enables teams to move faster with is not a DevOps-only or CI/CD-focused role. We are looking for engineers who build platform solutions not just configure pipelines.

Build and operate distributed systems using AWS managed services such as EKS ElasticCache and ML technologies like Ray over Kubernetes and NVIDIA Triton Inference internal tooling and automation frameworks to improve infrastructure reliability cost-efficiency and operational with engineering teams to define infrastructure architecture troubleshoot complex issues and drive production and manage Infrastructure as Code with Terraform ensuring repeatable secure and scalable or participate in incident response postmortems and continuous improvement cycles to reduce future risk.

3 years of experience in internet-facing backend production systems SRE or ML Operations focused roles on large scale distributed cloud infrastructurenProven expertise with AWS-managed infrastructure nFamiliarity with ML lifecycle and associated technologies such as NVIDIA Triton AnyScale Ray Apache Airflow programming skills in at least one of: Python Java Rust Go or similar languagesnHands-on experience with Linux systems and deep knowledge of its experience with Infrastructure as Code especially foundation in SRE concepts: Monitoring alerting observability Incident response and root cause analysis Error budgets SLAs/SLOs and system reliabilityn

Built tools or services that automate platform operations reduce toil or improve cost managing Kubernetes clusters at scale in production -on experience troubleshooting distributed systems under real-world communication skills and comfort collaborating across engineering infrastructure and product certifications or broad experience across multiple AWS services is a of modern GPU hardware architectures (such as NVIDIA H100 B200 or GB200 AWS Inferentia ) associated driversnUnderstanding of high-performance fabrics and network architecture power and thermal limits n

Required Experience:

IC


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile