Senior ML Engineer, Apple Ray, Apple Data Platform
Cupertino, CA - USA
Job Summary
Apple Ray integrates deeply with Apples data and ML ecosystem to provide a unified platform for building orchestrating and scaling complex ML and data pipelines. As a Software Engineer with ML background you will design distributed systems that support large-scale model training tuning and inference across heterogeneous compute environmentsfrom bare-metal GPU clusters to cloud-native will build features that enhance developer productivity for ML engineers improve resource efficiency and advance the performance and reliability of Apples ML workloads. Youll collaborate closely with ML practitioners to translate model and pipeline needs into robust platform capabilities while also improving the underlying distributed runtime and control role requires strong engineering fundamentals hands-on experience with ML systems and a passion for building scalable infrastructure.
Build scalable distributed systems and platform components using Ray that power Apples dataML APIs libraries and services that improve the efficiency and usability of large-scale ML training and inference performance and resource utilization across GPU/CPU clusters for ML workloads running at Apple with ML teams to understand model and pipeline needs and translate them into robust platform fault-tolerant orchestration mechanisms autoscaling strategies and runtime improvements for distributed ML complex issues across distributed systems and ML pipelines to ensure reliability and observability monitoring and debugging capabilities targeted at ML-centric distributed to architectural decisions and where appropriate upstream enhancements to Ray and related tools.
6 years building distributed systems high-scale backend services or compute background in ML workflows model training model serving or data pipeline in Python plus strong experience in a systems-level language (C Rust Go or Java).nExperience with ML frameworks such as PyTorch or TensorFlow and familiarity with GPU-based of parallelism strategies model scaling or distributed training with cluster orchestration (Kubernetes EKS GKE) or large-scale compute debugging skills across distributed and ML-centric runtime to work cross-functionally with ML engineers data engineers and infrastructure teams.n B.S. M.S. or Ph.D. in Computer Science Machine Learning or related technical fields or comparable software engineering experience.
Experience with distributed training frameworks (DeepSpeed Horovod FSDP ZeRO).nBackground in optimizing GPU workloads or performance with model orchestration systems or ML to open-source ML or distributed systems with large-scale data systems such as Spark Flink or similar.n
Required Experience:
Senior IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more