Senior Video SRE
San Diego, CA - USA
Job Summary
As a Senior Video Site Reliability Engineer at Apple you will be responsible for the reliability scalability and performance of our distributed applications that serve millions of users globally. You will build strong partnerships with application development teams sister SRE teams platform teams product teams as well as video encoding specialists to drive shared ownership of service reliability and maintain exceptional quality of experience for our day-to-day work will include embedding reliability practices into the development lifecycle building sophisticated monitoring and observability solutions and developing automation to reduce operational will own critical infrastructure components that video services depend on and use data-driven approaches to identify and eliminate single points of failure. You will lead reliability design reviews and define SLO frameworks that set the standard for service health. As part of the role you will participate in on-call rotations lead incident response efforts and drive post-incident reviews that result in meaningful reliability role offers the opportunity to work with complex JVM-based microservices and distributed systems technologies and influence architectural decisions that shape how Apple delivers video streaming content worldwide.
Bachelors degree in Computer Science Engineering or a related technical field or equivalent practical experience.n5 years of experience in Site Reliability Engineering DevOps or Systems Engineering with demonstrated senior-level ownership at scale including on-call/incident response post incident reviews and driving operational understanding of Linux fundamentals and networking principles with experience operating and debugging production in at least one programming language (Shell Python Go or similar) to reduce toil build SRE tooling and improve -on experience with cloud infrastructure and container troubleshooting and root-cause analysis skills across the full technology communicator who can collaborate with cross-functional partners to drive reliability outcomes.
Thorough understanding of distributed systems fundamentals failure modes and resilience patterns that prevent cascading record of building and continuously improving observability (metrics/logs/traces) alert quality and incident response processes for complex high-traffic -on performance optimization capacity planning and reliability engineering (load testing bottleneck analysis degradation strategies).nProven ability to build and operate Infrastructure as Code and CI/CD pipelines including safe deployment practices and change risk debugging and operating JVM-based applications in production (e.g. understanding of GC thread analysis heap profiling).nWorking knowledge of database systems key-value stores caching layers message queues and storage infrastructure at with video streaming technologies codecs protocols and media delivery infrastructure.
Required Experience:
Senior IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more