AI Research Scientist Multimodal Foundation Models Architecture, Pre-Training & Distillation

Apple


Job Location:

Seattle, WA - USA

Monthly Salary: Not Disclosed
Posted on: 13 hours ago
Vacancies: 1 Vacancy

Job Summary

The Multimodal Intelligence Team is building the next generation of foundation models for Apple experiences. We are looking for a research scientist to advance the architectures pre-training methods and distillation techniques that make highly capable multimodal models practical across the Apple research spans the full foundation-model lifecycle: model architecture pre-training objectives data mixtures optimization scaling distillation and evaluation. A defining challenge of our work is to develop models that combine broad intelligence with the memory latency energy and privacy requirements of on-device will have the opportunity to shape new research directions conduct ambitious experiments at scale and translate successful ideas into foundation-model technologies that can reach Apple products. Where appropriate this work may also lead to publications and the open sourcing of selected models research artifacts evaluations or tools.

In this role you will investigate fundamental questions about how multimodal foundation models should be designed trained and will develop and evaluate new model architectures pre-training objectives data strategies optimization methods and teacherstudent learning techniques. Your work will explore how capabilities developed in large foundation models can be effectively transferred to smaller more efficient models without treating distillation as an isolated downstream major focus of the role will be the co-development of frontier models and efficient models for Apple silicon and on-device intelligence. This includes designing architectures that distill effectively studying how teacher and student models should be trained together and developing distillation methods that preserve reasoning multimodal understanding instruction following and other important capabilities under constrained model than treating deployment constraints as an afterthought you will incorporate them into the research processfrom early architecture experiments and pre-training through distillation and final model may thrive in this role if you:nn* Want to invent new foundation-model architectures rather than only adapt existing models.n* Enjoy combining scientific ambition with real compute memory latency and energy constraints.n* Believe that small and efficient models can be a frontier research problem not merely a compression exercise.n* Are comfortable working across model research data systems and hardware boundaries.n* Care about translating research into private useful and deeply integrated intelligent experiences.n* Want your work to have both product impact and a presence in the broader research research directions include:nn* Novel dense recurrent state-space mixture-of-experts and hybrid foundation-model architectures.n* Multimodal pre-training across language images video audio and sensor-derived representations.n* Compute-optimal model and data scaling including data mixtures curricula tokenization and training objectives.n* Architecture and algorithm co-design for memory-efficient and energy-efficient inference on Apple silicon.n* Offline and on-policy distillation using teacher-generated data logits representations rationales and other supervision signals.

Hands-on experience designing implementing and running large-scale pre-training experiments for large language with LLM pre-training topics such as model architecture training objectives data mixtures tokenization curricula scaling and proficiency with modern deep learning frameworks such as PyTorch or JAX and distributed training evaluating pre-trained models across language understanding reasoning instruction following or multimodal understanding of transformer-based architectures and current approaches to efficient or scalable foundation-model degree or equivalent practical experience in machine learning computer science or a related technical field.

Experience contributing to major foundation-model pre-training efforts or leading architecture experiments that influenced a large training contributions in model architecture scaling laws multimodal pre-training optimization efficient attention mixture-of-experts state-space models or related with knowledge distillation including offline or off-policy distillation on-policy distillation self-distillation sequence-level distillation logic matching or representation designing teacherstudent training pipelines or transferring capabilities from large foundation models to smaller with multimodal models spanning language vision video audio or other sensor of inference efficiency memory hierarchy hardware accelerators or hardwaresoftware publication record influential open-source contributions or an equivalent record of applied research impact

Required Experience:

Staff IC

The Multimodal Intelligence Team is building the next generation of foundation models for Apple experiences. We are looking for a research scientist to advance the architectures pre-training methods and distillation techniques that make highly capable multimodal models practical across the Apple re...

About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile