AIML Distinguished Engineer, Foundation Model
Cupertino, CA - USA
Job Summary
engineered specifically for Apple silicon and for experiences that are privatenpersonal and deeply integrated into the OS. Behind that modeling work sits andemanding systems layer and the inference engine is at its inference engine is used by the foundation model team throughout the modelndevelopment lifecycle: generating and processing data and running rollouts forntraining powering large-scale model evaluation and serving as an LLM judgenthat scores and compares model outputs. These workloads are high throughputnbursty and tightly coupled to research iteration the speed efficiency andnreliability of the engine directly set the pace at which the team can train andnimprove a Distinguished Engineer you will own the technical strategy for thisninference engine and the broader systems that support it. You will partnernclosely with modeling and research teams to bring new capabilities into thendevelopment loop work across many internal teams with very differentnrequirements and lead a diverse set of engineers in turning an ambitiousnvision into shipped milestones. While inference is the initial focus you willnalso help shape adjacent areas training infrastructure data systems andnevaluation. If you are drawn to hard systems problems where the research andnthe infrastructure are inseparable this is the role.
Set and drive the technical vision and roadmap for the foundation modeln teams inference engine and the systems around it used for trainingn evaluation and LLM-as-judge deep work on inference performance efficiency and reliability:n throughput and latency optimization batching and scheduling quantizationn speculative decoding KV-cache management memory and compute efficiency andn hardware-aware inference systems that support a wide range of internal use cases n data generation and rollouts for training offline and large-scale evaluationn and judge/reward scoring across text image speech and multi-modal modelsn each with distinct throughput cost and quality your impact into adjacent systems areas training infrastructure datan pipelines and evaluation with many teams that depend on the engine translating their diversen needs into a coherent platform clear interfaces and a prioritized closely with ML researchers and modeling teams to co-design models andn systems and to bring state-of-the-art techniques from prototype into then development loop a diverse set of engineers across teams in setting direction andn executing against it; align stakeholders resolve technical trade-offs andn make the calls that keep large efforts prioritization and milestone delivery across competing demands balancingn near-term research needs against long-term platform and grow junior and senior engineers; establish engineeringn standards review designs and multiply the impact of the organization.
MS or PhD in Computer Science Machine Learning or related technical fieldn or equivalent industry experience.n15 years of experience building large-scale ML or distributed systems withn a track record of technical leadership and industry-wide or company-widen hands-on expertise in foundation model inference engines with a provenn record of improving performance efficiency and reliability at experience supporting a diverse set of foundation model inference usen cases each with different throughput latency cost and quality beyond inference the ability to contribute in adjacent systems areasn such as training infrastructure data systems or understanding of GPU/TPU/accelerator architecture distributed systems andn model optimization (quantization distillation compilation serving).nProficiency with ML frameworks such as JAX PyTorch and withn inference/serving experience leading a diverse set of engineers in setting vision andn driving execution including prioritization for milestone experience mentoring junior and senior experience partnering with ML researchers and modeling teams ton productionize research.
Experience building or leading inference systems for large language models andn multi-modal foundation models at with inference in training evaluation or reinforcement-learningn loops (e.g. large-scale rollouts offline eval or LLM-as-judge / rewardn scoring).nFamiliarity with Kubernetes Docker and cloud platforms (AWS GCP Azure)n and with distributed computing of defining technical strategy that shaped an organizations or then industrys direction.
Required Experience:
IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more