Enter a job title or keyword

Sr Manager Infrastructure, SRE, & AI Platforms Services Special Projects

Apple


Job Location:

Cupertino, CA - USA

Monthly Salary: Not provided by the employer
Posted: 26 August 2026 (2 days ago)
Application Deadline: 23 November 2026
Vacancies: 1 Vacancy

Job Summary

We are looking to hire a Senior Infrastructure SRE u0026 AI Platforms Manager to help set the long-term technical strategy organizational structure and operational roadmap for global mission-critical infrastructure platforms on the Services Special Projects team. nnThis position requires a rare blend of deep technical domain expertisespanning distributed systems Kubernetes and AI workload orchestrationand proven organizational leadership managing large globally distributed engineering teams.n

In this role you will be responsible for defining and building infrastructure strategy that balances continuous innovation with high reliability performance and cost efficiency. You will lead a growing multi-tiered team of engineers who are responsible for foundational platforms that power large-scale consumer and enterprise operational delivery you will establish standards for operational excellence Site Reliability Engineering (SRE) and capacity planning. You will be a key strategic partner translating complex business imperatives into scalable platform designs while cultivating a strong engineering culture focused on automation technical ownership accountability and continuous improvement.

Strategic Leadership u0026 Architecturenn Multi-Year Roadmap u0026 Strategy: Define and execute the long-term technical vision and capital investment strategy for global compute storage network observability and AI infrastructure.n Management u0026 Organizational Alignment: Partner with Leadership to align platform capabilities risk management capacity investments and architectural decisions with overarching business goals.n Technical Tradeoffs: Evaluate emerging infrastructure technologies and make strategic platform trade-off Compute u0026 Modern Infrastructure Platformsnn AI Infrastructure at Scale: Architect scale and optimize large-scale environments for training and inference resolving complex challenges in cluster design scheduling interconnect performance storage throughput and capacity planning.n Hybrid u0026 Multi-Cloud Compute: Oversee internal Kubernetes compute environments as well as managed public cloud platforms (AWS EKS GCP GKE) and large bare-metal footprints to provide seamless developer experiences.n Data u0026 Storage Platform Management: Direct the strategy and maintenance for distributed block/object storage alongside managed database and data streaming platforms (e.g. Cassandra FoundationDB Redis PostgreSQL MongoDB Kafka).n Networking Traffic u0026 Security: Ensure reliable global traffic management load balancing cloud networking architectures and enterprise security compliance across all Operational Excellence u0026 Engineering Culturenn Site Reliability Engineering (SRE): Cultivate a mature SRE culture focusing on high availability automated fault recovery telemetry logging metrics and rigorous post-incident analysis.n Global Team u0026 Leadership Development: Build mentor and lead a globally distributed organization comprising engineers managers and managers-of-managers across all levels (interns through senior principal staff).n Culture of Ownership u0026 Automation: Establish an environment characterized by strong technical ownership clear accountability continuous operational refinement and aggressive automation of manual processes.

MS Degree in Computer Science or related degree and 12 years of experience of progressive engineering leadership experience building scaling and operating mission-critical infrastructure platforms and global u0026 Leadership Scope: 6 years managing multi-layered engineering organizations (manager-of-managers) with a proven track record of hiring developing and retaining top-tier technical talent across global u0026 Distributed Compute Expertise: Demonstrated hands-on and architectural mastery of cloud-native infrastructure Kubernetes platform engineering and hybrid cloud operations (AWS GCP private data centers).nAccelerated Computing u0026 AI Infrastructure: Direct operational and architectural experience running large-scale systems for AI/ML training and inference workloads including utilization optimization scheduling and high-performance storage/ u0026 Production Operations: Deep background in Site Reliability Engineering (SRE) principles telemetry observability frameworks disaster recovery and managing 24/7 high-availability infrastructure at Communication: Exceptional ability to seamlessly bridge executive strategy and low-level technical trade-offscommunicating vision to executive stakeholders while driving detailed technical discussions with principal engineers.

Large-Scale Enterprise Provenance: Experience leading core infrastructure or foundational platform SRE for a global tier-1 technology organization operating at massive -Engine Database u0026 Data Infrastructure: Familiarity overseeing diverse open-source and proprietary storage/data ecosystems (e.g. Cassandra FoundationDB Kafka Redis PostgreSQL).nFinancial u0026 Capacity Governance: Proven competency managing large-scale infrastructure investments capital expenditures operational budgets capacity forecasting and cloud optimization strategies.

Required Experience:

Manager


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile