StaffSr. ML Infrastructure Engineer, Foundation Model Compute Infra
Cupertino, CA - USA
Job Summary
As a Senior/Staff Engineer on the Foundation Model Compute Infrastructure team you will design and build large-scale infrastructure that powers foundation model training fine-tuning evaluation and inference. You will develop model inference and fine-tuning services onboard and benchmark new accelerators and work closely with foundation model researchers and engineers to improve reliability performance scalability and developer productivity across Apples AI workloads.
Design build and evolve large-scale model serving and fine-tuning services for foundation model workloadsnDevelop reliable infrastructure for model deployment serving autoscaling traffic management job execution container orchestration and serving performance analysisnImprove the performance and usability of model serving and fine-tuning workloads by optimizing latency throughput availability accelerator utilization checkpoint loading compilation caching KV-cache-aware routing and workflows for launching monitoring debugging evaluating and deploying modelsnOnboard and Benchmark new accelerator technologies into Apples compute infrastructurenCollaborate with the Apple Foundation Model team to integrate technologies such as Pathways Ray and Beam or expose them as reliable and scalable servicesnMentor engineers and partner across teams to influence the technical direction of Apples foundation model compute infrastructure
5 years of industry experience building large-scale distributed systems or cloud infrastructurenExperience with distributed ML training or inference systemsnStrong programming skills in Python Go C or similar systems languagesnExperience with accelerator infrastructure such as TPU GPUnExperience with Kubernetes container orchestration or large-scale cluster management systemsnStrong communication and collaboration skills across engineering and research teamsnBachelors degree in Computer Science Engineering or related field
Experience building schedulers resource managers or orchestration systems for distributed workloadsnFamiliarity with frameworks such as JAX PyTorch TensorFlow Ray Pathways or vLLMnExperience operating large-scale multi-tenant infrastructure in cloud or hybrid environmentsnBackground in performance optimization fault tolerance or resource efficiency for large distributed systemsnStrong expertise in distributed systems scalability reliability and performance engineeringnExperience designing backend services or infrastructure platforms operating at production scalenMS or PhD in Computer Science Engineering or related field
Required Experience:
Senior IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more