Enter a job title or keyword

Software Engineer, Platform Reliability Engineering, AiDP

Apple


Job Location:

Sunnyvale, CA - USA

Monthly Salary: Not provided by the employer
Posted: 2 October 2026 (22 hours ago)
Application Deadline: 30 December 2026
Vacancies: 1 Vacancy

Job Summary

AI u0026 Data Platforms (AiDP) is ISu0026Ts engine for AI-powered innovation. The team brings together data application development and machine learning including generative AI along with data services and customer success functions to help ISu0026T build solutions more efficiently and streamline the adoption and embedding of generative AI across Applied Machine Learning team in AI and Data Platform organization is building the foundation for Apples enterprise-wide machine learning and data capabilities. Our Applied Machine Learning team designs builds and operates mission-critical platforms and services spanning ML GenAI inference and big dataenabling teams across the company to harness AI and analytics at scale. We tackle complex technical challenges in reliability performance and scalability across a diverse ecosystem of open source and cutting-edge technologies serving some of Apples most demanding workloads.n

Were seeking an experienced software engineer to join our Platform Reliability Engineering team and drive the design operation and optimization of large-scale distributed systems that power our GenAI ML and big data platforms. Youll leverage cutting-edge open source technologies in hybrid cloud environments to build resilient infrastructure that enables seamless inference data processing and machine learning this role youll own mission-critical platform components respond to production incidents and collaborate across teams to shape the future of our data and AI infrastructure.n

Design build and maintain scalable multi-tenant systems that support diverse workloads and technologies at enterprise scalenOwn the full lifecycle of infrastructure and platform projectsfrom architectural design and implementation through deployment monitoring and optimizationnOperate and optimize high-throughput mission-critical services to ensure reliability performance and cost-efficiencynParticipate in on-call rotations to respond to production incidents; diagnose root causes implement rapid fixes and drive post-incident improvementsnLead cross-functional collaboration with engineering teams to define requirements validate designs and deliver customer-impacting features and improvementsnProactively identify operational bottlenecks and systemic issues; implement preventive measures to reduce incident frequency and improve system resiliencenEstablish observability practices and continuously refine operational excellence standards across the platform

Bachelors degree in Computer Science Computer Engineering or equivalent professional experiencenProficiency in at least one systems programming language (Python Go Java or similar)nStrong expertise in distributed systems architecture with deep knowledge of reliability scalability and containerization principlesnHands-on experience with cloud platforms and data processing infrastructure (Kubernetes Spark Flink Ray Trino or equivalent technologies)n

7 years of experience in SRE DevOps or infrastructure engineering with demonstrated expertise managing distributed systems at in diagnosing and resolving complex production incidents and performance bottlenecks in large-scale distributed with open source codebases; ability to read understand and explain complex system implementationsnStrong understanding of system architecture and proven ability to collaborate effectively across engineering teamsnHands-on experience with big data technologies (Spark Flink Iceberg) and/or ML/AI platforms (Ray MLflow model serving infrastructure).nStrong foundational knowledge of Linux databases and security principlesnProactive mindset with demonstrated commitment to optimizing reliability and uptime for mission-critical servicesnExcellent written and verbal communication skills with ability to articulate technical concepts and strategies to both engineering teams and non-technical leadershipnDemonstrated track record of designing and operating systems at scalen

Required Experience:

IC


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile