Enter a job title or keyword

Staff ML Infrastructure Engineer

Apple


Job Location:

Cupertino, CA - USA

Monthly Salary: Not provided by the employer
Posted: 1 October 2026 (3 days ago)
Application Deadline: 29 December 2026
Vacancies: 1 Vacancy

Job Summary

Join a team at the forefront of ML infrastructure and generative AI where data and model workflows come together to enable the next generation of intelligent experiences on Apple products and services. We build robust systems that connect scalable data pipelines with advanced ML workflows accelerating the development of real-world AI applications. Our work spans the full ML lifecycle from experimentation to deployment and youll play a key role in shaping how AI models are built optimized and scaled. We develop a platform for ML data and features that powers advanced GenAI applications. This includes embeddings (generation evaluation ANN search multimodal support) AI Ops efficient inference and a modern feature platform designed to streamline experimentation and drive innovation. Were looking for engineers and researchers passionate about generative models data-centric ML and intelligent systems across diverse real-world use cases. With the autonomy to experiment the scale to make an impact and the support to take ideas from prototype to production youll work alongside a world-class team to build intelligent flexible systems that make ML development faster more reliable and more creative. n

The Apple AI Platform team gives Apples ML engineers and researchers the data systems and large-scale compute they need to build and ship models at Apples bar for quality and privacy. Our team owns the data layer that large-scale model training depends on: ingestion versioning lineage and governance on the way in and high-throughput data loading into the training fleet on the way out. As a Staff ML Infrastructure Engineer you will set the technical direction for that platform and own its hardest system-level problems the architecture other engineers and teams build on.

Own the architecture of the platform behind Apples largest model builds: define how ingestion immutable versioning lineage and governance work across structured unstructured and multimodal data at petabyte scale so every model run is reproducible from a versioned the technical direction for high-throughput data delivery to Apples largest GPU and TPU fleets: define the data access and loading architecture that keeps training compute-bound not I/ the hard system-level and format calls that the whole platform inherits columnar and lakehouse strategy the dataset abstraction spanning structured and multimodal data the shape of the SDK and core libraries backed by design and proof not just technical direction and influence across the platform and partner teams (data embeddings features research) and define the interfaces and contracts between the technical bar across the team: mentor senior engineers lead design reviews and be the escalation point for the problems no one else can with research and product leadership to shape the platform roadmap for next-generation workloads: foundation models multimodal data and retrieval-augmented efficiency reliability and automation across the data plane and control plane that power Apples ML fleet.

10 years of work experience in machine learning infrastructure distributed data systems or a related field.n10 years of experience building and shipping large-scale data or ML infrastructure and platforms in experience architecting and delivering large-scale distributed data or ML infrastructure that multiple teams or products depend on in track record of setting technical direction and driving it to delivery across teams not just within a single systems engineering: strong Python plus a systems language (Rust strongly preferred; C or Go acceptable) and hands-on performance engineering for I/O-bound workloads (Arrow zero-copy memory mapping async I/O high-throughput object storage).nDeep familiarity with columnar and lakehouse formats (Parquet Iceberg Delta or Lance) and the judgment to choose between them at working knowledge of the end-to-end ML workflow and how training and inference consume data enough to architect data systems that serve with modern ML and generative techniques (transformers diffusion retrieval-augmented generation fine-tuning) at the level needed to design for those ability to design highly available easy-to-use systems and to mentor and elevate the engineers around collaboration and communication with the ability to align multiple teams around a technical .S. M.S. or Ph.D. in Computer Science Computer Engineering or equivalent practical experience.

Experience defining data or ML platform architecture that was adopted across an experience with the data-loading and dataset-access layer of a modern ML framework (PyTorch JAX or TensorFlow).nDistributed data-loading frameworks for ML: Ray Data NVIDIA DALI WebDataset or Mosaic feeding data to GPU or TPU fleets at scale and keeping them lineage and governance systems: DataHub OpenLineage Unity Catalog or to or operational experience with Spark Daft Polars or DuckDB and orchestration (Docker Kubernetes).

Required Experience:

Staff IC


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile