SRE Manager, ML Operations
New York City, NY - USA
Job Summary
We are looking for a senior engineering leader to manage and grow our Site Reliability Engineering team with a focus on ML Operations. This team owns the reliability performance and scalability of the Ad Serving infrastructure that serves as the critical front door of Apple Ads operating at one of the largest scales in the is a high-impact leadership role where you will shape the future of how we build run and evolve our ML Platforms and Services globally. You will bring deep technical expertise while staying anchored to business and product goals and you will cultivate a team culture defined by operational excellence innovation and continuous improvement.n
Lead and scale SRE teams responsible for the reliability performance and availability of ML Platforms and ServicesnGrow and develop your engineers through mentorship clear goal-setting and meaningful career developmentnDefine a compelling team vision and drive execution toward high-quality measurable outcomesnChampion reliability engineering best practices including SLOs/SLAs error budgets observability incident management and fault analysisnFoster a culture of engineering excellence -- encouraging innovation knowledge sharing and continuous improvementnPartner with staff engineers and technical leadership on architecture decisions platform strategy and long-term roadmapnCollaborate cross-functionally with Product ML Platform Ads Serving and Data Science teams to deliver complex high-impact initiatives
10 years of experience with large-scale distributed systemsn5 years of experience in an engineering leadership role ideally managing SRE or Production Engineering teamsnProven track record of building and leading high-performing engineering teamsnStrong grasp of core operating system principles networking fundamentals and systems managementnDeep understanding of SRE principles: monitoring alerting error budgets fault analysis capacity planning and incident responsenExcellent problem-solving communication and decision-making skills
Bachelors or Masters degree in Computer Science or a related fieldnExperience managing and optimizing GPU-based clusters in production environmentsnExperience building and operating large-scale ML systems or ML infrastructure at scalenHands-on experience managing cloud infrastructure particularly AWSnFamiliarity with the digital advertising ecosystem and its technical demandsnDemonstrated ability to influence and partner across Product Data Science and Platform Engineering organizations
Required Experience:
Manager
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more