The Apple Services Engineering team (ASE) is one of the most exciting examples of Apples long-held passion for combining art and technology. These are the people who power the App Store Apple TV Apple Music Apple Podcasts and Apple Books at extensive scale meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 ASE the Apple Data Platform SRE team keeps a massive multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure automation and customer success running incident response providing hands-on support to internal teams and partnering with developers to make cutting-edge services like Spark Flink Airflow Ray Notebooks and LLM-based agent platforms reliable at scale.n
This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple while specialising in an area thats shaping the future of how Apple builds and operates AI. As an SRE on Apple Data Platform youll operate and support the teams full portfolio from big data pipelines to multi-cloud infrastructure and grow into the teams go-to expert for ML/AI platform services including Ray training and serving LangGraph agent deployments RAG architectures embeddings platforms and vector store platforms. You wont be building the models yourself but youll be the infrastructure backbone behind the teams who do keeping their services pipelines and platforms running flawlessly in production so they can focus on looking for a self-motivated engineer who thrives on ownership someone who wants a set of services to call their own the autonomy to drive their reliability roadmap and the collaborative instinct to keep that work aligned with the teams broader direction. If you love solving hard operational problems enjoy being the trusted expert customers turn to and want a front-row seat to Apples ML/AI infrastructure evolution this role offers real room to grow your scope and impact over time.
Operate monitor and triage production and non-production environments across the ADP portfolio data processing ML/AI and multi-cloud in a rotating on-call schedule across supported services including occasional weekday and weekend the operational health of ML/AI platform services as SME driving reliability support and customer guidance for Ray LangGraph agents RAG pipelines embeddings and vector store Slack-based support to internal customers; screen triage and resolve service related with dev teams across time zones to onboard new services understanding architecture then designing monitoring alerting and dashboards (Prometheus Grafana Splunk).nBuild automation and self-healing tooling that reduces manual toil and scales the teams operational escalate and resolve production issues to protect platform reliability and customer with SRE and dev partner teams engineering and program management to align execution with team and org goals.n
Minimum Qualificationsnnn* Bachelors Degree in Computer Science an engineering-related field or equivalent related experience.n1-4 years in a Site Reliability Engineering DevOps or Infrastructure-focused in Python; working knowledge of Golang a with Kubernetes and at least one major cloud provider (AWS or GCP).nExposure to operating or supporting ML pipelines model-serving infrastructure or LLM-based systems in communication skills and composure under pressure during grounding in SRE principles with prior on-call or production-support experience.n
Hands-on experience operating or supporting Ray (training/serving) LangGraph or similar agent orchestration frameworks RAG architectures embeddings platforms or vector store with MCP-based tooling and ML lifecycle/dataset management with S3 and cloud storage/networking with observability tooling: Prometheus Grafana Splunk knowledge of CI/CD pipelines and deployment understanding of one or more Big Data technologies (Spark Flink Airflow Trino Notebooks).nA track record of automating manual operations through scripting or curiosity and a drive to keep learning for yourself your team and the org.
Required Experience:
IC
The Apple Services Engineering team (ASE) is one of the most exciting examples of Apples long-held passion for combining art and technology. These are the people who power the App Store Apple TV Apple Music Apple Podcasts and Apple Books at extensive scale meeting high expectations to deliver a hug...
The Apple Services Engineering team (ASE) is one of the most exciting examples of Apples long-held passion for combining art and technology. These are the people who power the App Store Apple TV Apple Music Apple Podcasts and Apple Books at extensive scale meeting high expectations to deliver a huge variety of entertainment in over 35 languages to more than 150 ASE the Apple Data Platform SRE team keeps a massive multi-cloud platform running for thousands of internal engineers building the next generation of data and AI products at Apple. We sit at the intersection of infrastructure automation and customer success running incident response providing hands-on support to internal teams and partnering with developers to make cutting-edge services like Spark Flink Airflow Ray Notebooks and LLM-based agent platforms reliable at scale.n
This is a rare opportunity to build deep expertise across one of the most technically diverse platforms at Apple while specialising in an area thats shaping the future of how Apple builds and operates AI. As an SRE on Apple Data Platform youll operate and support the teams full portfolio from big data pipelines to multi-cloud infrastructure and grow into the teams go-to expert for ML/AI platform services including Ray training and serving LangGraph agent deployments RAG architectures embeddings platforms and vector store platforms. You wont be building the models yourself but youll be the infrastructure backbone behind the teams who do keeping their services pipelines and platforms running flawlessly in production so they can focus on looking for a self-motivated engineer who thrives on ownership someone who wants a set of services to call their own the autonomy to drive their reliability roadmap and the collaborative instinct to keep that work aligned with the teams broader direction. If you love solving hard operational problems enjoy being the trusted expert customers turn to and want a front-row seat to Apples ML/AI infrastructure evolution this role offers real room to grow your scope and impact over time.
Operate monitor and triage production and non-production environments across the ADP portfolio data processing ML/AI and multi-cloud in a rotating on-call schedule across supported services including occasional weekday and weekend the operational health of ML/AI platform services as SME driving reliability support and customer guidance for Ray LangGraph agents RAG pipelines embeddings and vector store Slack-based support to internal customers; screen triage and resolve service related with dev teams across time zones to onboard new services understanding architecture then designing monitoring alerting and dashboards (Prometheus Grafana Splunk).nBuild automation and self-healing tooling that reduces manual toil and scales the teams operational escalate and resolve production issues to protect platform reliability and customer with SRE and dev partner teams engineering and program management to align execution with team and org goals.n
Minimum Qualificationsnnn* Bachelors Degree in Computer Science an engineering-related field or equivalent related experience.n1-4 years in a Site Reliability Engineering DevOps or Infrastructure-focused in Python; working knowledge of Golang a with Kubernetes and at least one major cloud provider (AWS or GCP).nExposure to operating or supporting ML pipelines model-serving infrastructure or LLM-based systems in communication skills and composure under pressure during grounding in SRE principles with prior on-call or production-support experience.n
Hands-on experience operating or supporting Ray (training/serving) LangGraph or similar agent orchestration frameworks RAG architectures embeddings platforms or vector store with MCP-based tooling and ML lifecycle/dataset management with S3 and cloud storage/networking with observability tooling: Prometheus Grafana Splunk knowledge of CI/CD pipelines and deployment understanding of one or more Big Data technologies (Spark Flink Airflow Trino Notebooks).nA track record of automating manual operations through scripting or curiosity and a drive to keep learning for yourself your team and the org.
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar
... View more