SRE Software Engineer
Austin, TX - USA
Job Summary
The ASE Compute team is looking for a Site Reliability Engineer to deploy and manage a large Kubernetes platform that Apples services run on partnering with engineering teams across the company to solve complex problems using both open-source and in-house tooling. You will contribute to the development of our controllers and namespace management infrastructure working alongside senior engineers to strengthen the reliability of our Kubernetes services. You will learn to write well-tested code participate in design reviews and gradually take ownership of features. Youll have the opportunity to engage with the upstream community gain hands-on experience with production-scale systems and build the technical foundation to support service teams across Apple. The role also offers room to build AI-assisted tooling that accelerates triage operational workflows and infrastructure automation for the whole team.
Deploy configure and maintain large-scale multi-tenant Kubernetes environmentsnWrite and maintain operational tooling to improve reliability and reduce manual interventionnImplement and maintain reliability standards for the platform: SLOs error budgets alerting philosophy upgrade and rollout strategy and the run-books that follow from to CI/CD pipelines revision control workflows and configuration management practicesnTake on-call troubleshoot production issues and follow up on post-incident action itemsnHelp enforce security best practices OS hardening and compliance standards across the fleet
Hands-on experience in Linux systems administration and containerization with enterprise distributions such as RHEL Oracle Linux or CentOSnProficiency in Python or Go for scripting and toolingnSolid understanding of Linux fundamentals: file systems process management user and group administration and package managementWorking knowledge of networking concepts including TCP/IP DNS DHCP and basic firewall configurationnExperience with version control systems such as Git and configuration management (Puppet Ansible or equivalent)nStrong written and verbal communication skills
Site Reliability Engineering DevOps or Infrastructure focused experiencenExperience with third-party cloud platforms (AWS GCP or Azure)nExperience with containerization and orchestration technologies such as Docker or KubernetesnFamiliarity with bare-metal provisioning and lifecycle management at datacenter scalenUnderstanding of cloud-native observability (Prometheus Thanos Splunk or similar)nFamiliarity with CI/CD pipelines and DevOps practicesnKnowledge of OS security hardening encryption and regulatory compliance frameworks
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more