Enter a job title or keyword

AIML Machine Learning Research Lead, RL Agents, MLR

Apple


Job Location:

Cupertino, CA - USA

Monthly Salary: Not provided by the employer
Posted: 11 September 2026 (16 hours ago)
Application Deadline: 9 December 2026
Vacancies: 1 Vacancy

Job Summary

We are looking for a hands-on research lead to drive our work on reinforcement learning and post-training for agentic AI and to manage a small team of senior researchers working on related problems in RL agentic tool-calling synthetic environment generation model scaling and multimodal action models. You will help set direction for how we develop infrastructure training runtime and evaluation procedures for interactive agents tool calling coding computer use and long-horizon role sits inside a research organization pursuing first-principles approaches to core AI problems: generative foundation models across modalities (text images graphs scientific and engineering data) vision-language modeling and implicit world modeling self-supervised learning and search and evolutionary methods for optimizing both agents and the environments they learn in. A distinctive part of our agenda is designing methods that fit Apples deployment reality on-device and hybrid (device plus private cloud) execution co-designed with current and future hardware and that take advantage of what this ecosystem uniquely enables such as deeply personalized long-context agentic experiences. We aim for both field-changing research and direct impact on Apple products and internal engineering is a research group first. Management here is about spreading the load of running a team not stepping away from the work everyone including leads stays hands-on. We support continued engagement with the academic community: publishing conference service student collaboration and internships.

* Lead research on RL and post-training for agentic capabilities: reward preference optimization and verifier design training recipes and evaluation for tool calling coding and multi-step interactive * Build and own synthetic data and task-generation pipelines generating diverse verifiable tasks and environments along with the interactive environments and benchmarks that go with them and the curricula that turn them into capable * Drive codebases and infrastructure for the core RL research effort and help engage partner teams to use and co-develop the * Manage and mentor a small team (34) of senior researchers and research engineers with distinct specialties shaping a shared research direction while protecting room for bottom-up idea-driven * Stay hands-on: run experiments write code and contribute directly to the teams most important technical * Connect post-training research to efficiency and deployment: what works under on-device and hybrid compute constraints and how method design interacts with * Collaborate across the organization on adjacent directions including methods for environment and agent co-optimization self-improvement world models used as planners or policies and personalized long-context * Publish in top venues and engage with the broader research community.

PhD in machine learning or a related field or equivalent research experiencenn7-10 years of research experience beyond PhD in industry or as an academic research leadnnStrong track record in RL and/or post-training of large models demonstrated through publications open-source contributions or shipped systemsnnLeadership experience: setting and defending a research direction over multiple years and directing others work through direct reports PhD students postdocs or sustained project teams. Formal management experience is welcome but not requirednnExperience owning ML infrastructure frameworks and codebases including open-source research frameworks or environment suites others build onnnStrong engineering skills; comfortable working hands-on in large training codebases

Experience taking research from idea to product or production impactnnFamiliarity with efficiency-aware modeling: small models mixture-of-experts quantization distillation inference-cost constraints or hardware-aware method designnnInterest or background in open-endedness evolutionary computation curriculum or environment design multi-agent systems or self-improving systemsnnPrincipled or theoretical grounding in RL representation exploration or optimization views of policy learning alongside strong empirical worknnBreadth across core machine learning generative models self-supervised learning pre-training and perspective on the fields longer arcs not only its most recent methodsnnExperience growing other researchers and managing researchers and engineers with heterogeneous specialties and synthesizing their work toward a common goalnnExperience owning a large RL or post-training codebase

Required Experience:

Senior IC


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile