Enter a job title or keyword

Software Engineer Generative Data Platform, Evaluation

Apple


Job Location:

San Francisco, CA - USA

Monthly Salary: Not provided by the employer
Posted: 2 October 2026 (18 hours ago)
Application Deadline: 30 December 2026
Vacancies: 1 Vacancy

Job Summary

Apples machine learning features are only as good as the data behind them and our team builds the platform that lets teams create data that fills the gaps real-world collection cant provides privacy-preserving ways to train and evaluate our models safely and helps us ensure our products have seen diverse inputs to generalize properly across the many parts of the world where we the AI u0026 Machine Learning (AIML) organization our team builds the self-service tools and platform that turn ideas text and images into large-scale generated datasets. We dont produce the datasets ourselves: ML engineers data scientists designers feature teams ML data operations teams and Safety and Responsible AI teams across Apple use our platform to generate their own. That data is only valuable if it matches the real world closely enough to train on and to evaluate against proving our models and features meet Apples quality bar before they reach customers. As a software engineer on the platform youll sit between that technology and the people who depend on it: understanding what they want to create building the features that make it possible and helping them get the most from generative AI. Its a hands-on engineering role for someone who enjoys the people side as much as the code.

Youll spend your days moving between engineering and partnership. Synthetic data here does far more than fill gaps: it lets teams build privacy-preserving digital humans represent locales and domains that real-world collection underserves and experiment in days instead of waiting on slow costly data collection. It also gives Safety and Responsible AI red teams the data they need to stress-test models against misuse and edge cases. One day you might be integrating a new image or video generation model and tuning how it runs on GPUs; the next you might be helping a design team get their first dataset running on the platform or working with a feature team on model ablations that measure what synthetic data adds to training and evaluation. nnYoull help teams understand where synthetic data differs from the real data their models see whether thats how images look or subtler differences in metadata distributions and labels and then build what closes those gaps. Your core partners are ML engineers and data scientists but youll also work with designers artists and others who are newer to machine learning and help make these systems understandable and approachable for them. Youll own features end to end from shaping the request with partner teams to validating the result on production infrastructure. nnOur team builds with AI coding agents every day and youll help shape how we use them well. You dont need a research background in generative models; youll learn the models on the job. What matters most is strong engineering judgment and the ability to bring people along.

Partner with ML engineers data scientists and designers to turn their needs into platform teams as they build their own datasets on the platform and help design model ablations that measure how synthetic data affects training and how ML vision and generative models behave to people new to them including cost quality speed and the domain gap between synthetic and real data covering visual properties and non-visual ones such as metadata value distributions and with feature teams to find where synthetic data falls short for their use case and build platform capabilities to close those teams judge generated output understand why results vary and get closer to what they new image and video generation models into a multi-stage production and debug GPU workloads to keep generation reliable and cost-effective.n

Bachelors degree in Computer Science Computer Engineering or a related field or 3 years of equivalent work 3 years of experience building and operating production software in debugging distributed systems or data pipelines in production using logs metrics and task state to find root deploying machine learning models on GPU infrastructure including dependency management and GPU memory designing input validation and configuration contracts for systems where late failures are ability to explain machine learning concepts and tradeoffs to non-technical partners such as designers or product hands-on use of agentic AI coding tools in day-to-day software development including judging where they are effective and where their output needs human verification.n

Experience working directly with partner teams or internal customers to define and ship measuring the domain gap between synthetic and real data including visual statistics and non-visual properties such as metadata and label with generative image video or large language model APIs including structured with image and video file formats and metadata with batch compute platforms or job schedulers.n

Required Experience:

IC


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile