ML Infrastructure Engineer ML Compute Capacity
Santa Clara County, CA - USA
Job Summary
As an engineer on the ML Compute Capacity team you will design build and operate the production systems that ensure compute resources are optimally distributed throughout the company. Youll work across the stack from data pipelines and backend services to APIs and interactive frontends developing telemetry systems optimization algorithms policies and intuitive tools for managing demand and improving efficiency across Apples largest accelerator fleet. Our small nimble team works in a high-autonomy fast-paced environment and were passionate about digging into data patterns laying out the performance characteristics of an entire distributed system and knowledge sharing. If the opportunity to own and operate services that scale stay highly available and just work excites you then please reach out to us!
Build and operate demand and capacity planning systemsnBuild data pipelines and telemetry systems that ingest normalize and serve fleet-wide utilization and cost data across multi-tenant and heterogeneous fleetsnDevelop observability infrastructure monitoring alerting and dashboards that surfaces real-time fleet health and efficiency signalsnDrive innovation in forecasting optimization and supply chain management tooling that works at scalenBuild end-to-end tooling from data models and APIs to interactive dashboards that distills complex data into actionable insights for leadershipnBuild self-service platforms with well-defined schema contracts and APIs enabling ML teams infrastructure engineers and finance to balance usability utilization and costsnEngage cross-functionally with finance analysts supply chain managers data center operations compute infrastructure engineers and morenSupport the team through code reviews and knowledge sharing
7 years of experience in relevant areasnExperience with machine learning infrastructure on GPUs or TPUsnProficiency in Python and/or Go for production backend and data engineering worknExperience building data pipelines and crafting robust queries over large-scale multi-source data (e.g. Trino PostgreSQL Elasticsearch)nExperience with observability tools (e.g. Prometheus Grafana) or equivalent monitoring systemsnExcellent problem-framing and problem-solving skillsnStrong CS fundamentalsnBachelors degree or higher in Engineering Mathematics Economics or a related quantitative field
Experience operating Kubernetes at production scale including scheduling resource management and cluster debuggingnExperience with modern web frameworks like ReactnFamiliarity with accelerator utilization patterns across ML training and inferencenStrong interest with capacity planning cost attribution or FinOps systems
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more