Site Reliability Engineer Kafka

Apple


Job Location:

Seattle, WA - USA

Monthly Salary: Not Disclosed
Posted on: 2 hours ago
Vacancies: 1 Vacancy

Job Summary

The Apple Service Engineering Data Streaming SRE team is looking for Site Reliability Engineers with experience developing processes tools and automation for managing distributed systems in production environments. Our SRE team combines software engineering systems engineering and Devops practices to build and run large-scale massively distributed fault-tolerant systems. Our software ensures that Apples services are reliable scalable and secure and we leverage both open-source and homegrown technologies to provide managed data infrastructure services. You will help build next-generation Kafka infrastructure and platform services collaborating cross-functionally with various ASE teamsfrom store and commerce to search and recommendations. Youll create platforms that can rapidly scale to serve data with very low latencies. You should be someone who isnt afraid to question assumptions thrives as a collaborative partner under tight deadlines and tackles complex problems with elegant technical solutions.

The Data Service SRE team develops applications and tooling that are safe reliable scalable and fast. This work requires an innovative spirit and an extraordinary degree of care and difficulty in engineering. Team members contribute to all major components of Kafka deployment infrastructure including maintenance automation control plane enhancements monitoring and alerting tooling/dashboards advanced deployment architecture focused on safety stability performance and scaling. nnCome join us at Apple Services Engineering and help us deliver services and applications that are fluid and responsive. You will collaborate with engineers from across Apple to define the metrics set targets uncover optimization opportunities and ship a service that will delight our customers. This role is for engineers who enjoy deep technical engineering that spans large cross-organizational projects. Your openness to learning and implementing new technologies will contribute to the continuous evolution of our organization. Good ideas are valued and rewarded.

Understanding of core SRE concepts - Monitoring Alerting Incident managementnDeep and wide performance engineering (design concepts profile-guided optimization)nService lifecycle mangement across bare metal and virtualized (EC2) kubernetes platformsnPrepare alert handling procedures run-books and collaborate with other SRE team communication and a high degree of customer focus when engaging with internal platform customersnAs a distributed team ability to work optimally with colleagues based in other locations is essentialnPrior experience with development or maintenance of Kafka infrastructure or similar data service is highly recommended

5 or more years of experience in support of internet-facing production services and distributed systems via deployments On Call and Incident Management.n5 or more years of experience running large scale infrastructure with a heavy reliance on automation toolingn5 or more years of experience troubleshooting and performance deep dive analysisnReal operational experience managing services at scale on KubernetesnProficient in one or more of the following programming languages: Java Go (golang) PythonnOperational experience deploying in and running on Datacenter and Cloud architectures (networking topologies host placement strategies and failure modes); design of multi-datacenter systems; failure domains; and wide-area motivated inquisitive with an aptitude to learn new technologies quickly and expertise developing and troubleshooting distributed systems and database storage developing critical internet services and/or platform with AWS GCP and IaC such as Terraform

Experience managing messaging services such as Kafka or other Data servicesnProficient in Java Go (golang) u0026 Python

Required Experience:

IC

The Apple Service Engineering Data Streaming SRE team is looking for Site Reliability Engineers with experience developing processes tools and automation for managing distributed systems in production environments. Our SRE team combines software engineering systems engineering and Devops practices ...

About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile