AIML Sr Machine Learning Engineering Manager, Evaluation
Cupertino, CA - USA
Job Summary
As a Senior Machine Learning Engineering Manager in AIML Evaluation you will lead the technical strategy and execution for agent evaluation and automatic optimization. You will own systems that evaluate foundation models and agents diagnose failure modes and use those signals to drive automated prompt context tool rubric and agent-harness improvements. You will also help establish the interfaces between evaluation and post-training so that high-value failures can be converted into targeted data environments reward signals and measurable model is a hands-on leadership role. You will prototype new approaches participate in architecture and code reviews design experiments and help your team translate emerging research into scalable evaluation and optimization pipelines. You will partner closely with Apple Foundation Models product engineering teams and other AIML groups to build an evaluation flywheel that connects real product behavior with model and agent refinement. You will also work across the organization to advance synthetic data generation for both evaluation and post-training with strong attention to data quality representativeness privacy and reproducibility.
Architects and builds scalable evaluation systems for foundation models and agents including benchmarks LLM-based evaluators simulation environments trajectory analysis and regression an end-to-end evaluation flywheel with Apple Foundation Models and product teams that connects observed failures to diagnosis targeted refinement post-training and measurable quality mentors and grows a small team of machine learning engineers while remaining deeply involved in technical design experimentation implementation and the technical strategy and roadmap for automatic prompt context tool rubric and agent-harness optimization for agentic development and model methods that convert evaluation findings into actionable model-improvement signals including targeted datasets synthetic trajectories reward or preference signals and optimization across AIML to design and scale synthetic data generation pipelines for evaluation and and adapts recent research in LLM and agent evaluation automatic optimization LLM-as-judge reward modeling test-time search and post-training to production-quality workflows.
8 years of professional experience in machine learning applied research or software engineering including experience building production ML systems or large-scale experimentation platforms.n3 years of technical leadership experience including direct people management of machine learning or software engineers and a demonstrated ability to mentor and grow strong technical or PhD in Computer Science Machine Learning Artificial Intelligence or a related technical hands-on programming and software engineering skills particularly in Python with experience building reliable ML pipelines using modern machine learning or deep learning experience with large language models or agentic systems including evaluation of multi-turn behavior tool use planning reasoning or other action-taking building automated evaluation methods such as LLM-based judges rubrics reward models simulation-based evaluation or scalable benchmark with at least one model or agent refinement area such as automatic prompt or context optimization post-training preference optimization reinforcement learning or agent-harness communication and collaboration skills with demonstrated ability to align research engineering and product teams around ambiguous technical problems.
Track record of applying recent machine learning research to production systems or high-impact product with automatic prompt or context optimization agent-search methods evaluator optimization or multi-objective optimization for agentic generating and evaluating synthetic datasets tool-use trajectories or multi-turn agent interactions including methods for filtering deduplication diversity and quality designing evaluation systems that combine offline benchmarks simulation human evaluation and product- or usage-derived with privacy-preserving or on-device machine learning and ability to influence technical strategy across organizational boundaries and communicate complex model-quality tradeoffs to senior technical leaders.
Required Experience:
Manager
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more