Site Reliability Engineer (SRE), Observability, London
London, KY - USA
Job Summary
Observability infrastructure is BIG. Operating at our scale across multiple geographically dispersed data centers and servicing hundreds of millions of users presents unique challenges. As an SRE at Apple youll need to solve these problems using data teamwork and your own expertise. SREs at Apple own the full infrastructure stack; from device driver performance debugging to content delivery network traffic management our responsibilities are both broad and runs its systems on Linux and Kubernetes. We run a mix of open source vendor licensed and internally developed tools to perform functions such as system configuration management provisioning software deployment logging and monitoring. Youll learn these tools and have opportunities to improve them. Our team is collaborative; we work closely with the development teams we support to deliver the best results for Apple. We think critically and strive to balance the best solution with the need to get things done for each engineering challenge we face. Good ideas are heard and results are SRE is a small team with huge scale. We serve as central Telemetry and Logging for the Services organisation which powers most of the customer-facing Apple services like iCloud.
Deploy support and monitor new and existing services platforms and application stacks across production and non-production with our development and product partners to understand requirements then design and implement resilient scalable infrastructure architect author and deliver software that improves the availability scalability and security of Apples observability and run systems infrastructure and applications through automation รข provisioning configuration deployment and scale testing to measure tune and optimise system performance; contribute to capacity planning and disaster-recovery and integrate new technologies to improve system reliability security and on code infrastructure and design reviews and drive process in an on-call rotation providing hands-on technical expertise during service-impacting events.
Strong sense of ownership and integrity demonstrated through clear communication and collaborationnExperience in managing and scaling distributed systems in a public private or hybrid cloud environmentnThe ability to design author and release code in languages like (but not limited to) Go or PythonnAcute drive to automate manual operations and to improve them through repeated iterationnUnderstanding of the Linux Operating System standard networking protocols and components
Hands-on experience managing large numbers of diverse systems with configuration management or software delivery platforms (such as Puppet and Spinnaker)nExperience with deploying supporting and monitoring new and existing services platforms and application stacksnExperience with scale testing disaster recovery and capacity planningnFamiliarity with microservices architecture and container orchestration with Kubernetes
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more