ASE Compute Senior SRE Software Engineer
San Francisco, CA - USA
Job Summary
The ASE Compute team is looking for a senior SRE software engineer to own the technical direction of the Kubernetes internals that Apples services run on. You will set the architecture for our controllers and namespace management infrastructure strengthen the reliability of our Kubernetes services and engage with the upstream community to drive Apples requirements. You will write the hardest parts yourself and raise what the rest of the team can build through design review mentoring and the tools you build. Service teams across Apple will come to you as the technical authority on what the platform can do and where it is going. The role also offers room to build AI-assisted tooling that accelerates triage operational workflows and infrastructure automation for the whole team. nn
nOwn the architecture and technical direction for our controllers and namespace management infrastructure from design through production operation at global and fix the reliability problems that only appear past the point where upstream defaults and community guidance stop working build each fix into the platform so the same class of problem does not come back and automate the operational load that cannot be designed the reliability standards for the platform: SLOs error budgets alerting philosophy upgrade and rollout strategy and the runbooks that follow from on-call lead incident response for the hardest production issues and see the post-incident follow-up work through so the same incident does not happen Apple in the upstream Kubernetes community driving our requirements through design proposals and code in the relevant the engineers around you through design review code review and direct cross-team technical efforts through ambiguity and advise partner service teams and their leadership on platform capabilities and tradeoffs shaping both their designs and the Compute roadmap.n
Bachelors Degree in Computer Science an engineering-related field or equivalent related experiencen8 years in a Site Reliability Engineering DevOps or Infrastructure focused rolenDeep experience operating large-scale multi-tenant Kubernetes environments in productionnStrong systems background comfortable troubleshooting across the full stack (network OS container runtime application)nExpert-level Go with a track record of shipping and owning controllers or operators that other teams depend with configuration management at scale (Puppet Ansible or equivalent)nDemonstrated ability to drive cross-functional initiatives to completionnStrong written and verbal communication skills
Experience with third-party cloud platforms (AWS GCP or Azure)nFamiliarity with bare-metal provisioning and lifecycle management at datacenter scalenUnderstanding of cloud-native observability (Prometheus Thanos Splunk or similar)nExperience running infrastructure as an internal managed service with defined SLAs
Required Experience:
Senior IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more