Enter a job title or keyword

Multimodal Machine Learning Researcher

Apple


Job Location:

Cupertino, CA - USA

Monthly Salary: Not provided by the employer
Posted: 27 August 2026 (2 days ago)
Application Deadline: 24 November 2026
Vacancies: 1 Vacancy

Job Summary

Do you believe generative models can transform creative workflows and smart assistants used by billions Do you believe it can fundamentally shift how people interact with devices and communicate Our Scene Understanding strives to turn cutting edge research into compelling user experiences that realize all these goals and more working on Apple Intelligence technologies such as Image Playground Genmoji Generative Memories Semantic Search and many more. nnWe are looking for senior technical leaders experienced in architecting and deploying production scale multimodal ML. An ideal candidate has the ability to lead diverse cross functional efforts ranging from ML modeling prototyping validation and private learning. Solid ML fundamentals and an ability to place research contributions with respect to state of the art would be an essential part of the role. Experience with training and adapting large language models would be an important need. We are the Intelligence System Experience (ISE) team within Apples software organization. nnThe team works at the intersection between multimodal machine learning and system experiences. For example experiences like Spotlight Search Photos Memories Generative Playgrounds Stickers Smart wallpapers etc are all areas that the team has had a significant part in delivering through ML core technologies. These experiences that our users enjoy are backed by production ML workflows which our team works to scale through distributed training. Additionally our team also focuses on approaches to optimizing and adapting LLMs to best suit on-device user experiences. nnSELECTED REFERENCES TO OUR TEAMS WORK: - ( - ( - ( - ( are looking for a candidate with a proven track record in applied ML research. Responsibilities in the role will include training large scale multimodal (2D/3D vision-language) models on distributed backends deployment of compact neural architectures efficiently on device and learning policies that can be personalized to the user in a privacy preserving manner. Ensuring quality in the wild with an emphasis on fairness and model robustness would constitute an important part of the role. You will be interacting very closely with a variety of ML researchers software engineers hardware u0026 design teams cross functionally. The primary responsibilities of the role would center on enriching multimodal capabilities of large language models. The user experience initiative would focus on aligning image/video content to the space of LMs for visual actions u0026 multi-turn interactions.

M.S. or PhD in Computer Science or a related field such as Electrical Engineering Robotics Statistics Applied Mathematics or equivalent on experience training LLMs/adapting pre-trained LLMs for downstream tasks u0026 alignmentnModeling experience at the intersection of NLP and visionnProficiency in ML toolkit of choice e.g. PyTorchnStrong programming skills in Python

Familiarity with distributed trainingnStrong programming skills in C/C or ObjCn

Required Experience:

IC


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile