Enter a job title or keyword

Machine Learning Engineer Agentic AI Evaluation Frameworks

Apple


Job Location:

Cupertino, CA - USA

Monthly Salary: Not provided by the employer
Posted: 11 September 2026 (Yesterday)
Application Deadline: 9 December 2026
Vacancies: 1 Vacancy

Job Summary

Imagine what you could do here. At Apple great ideas have a way of becoming extraordinary products services and customer experiences very quickly. Bring passion and dedication to your work and theres no telling what you could Channel Sales AI Product Engineering team is looking for a Machine Learning Evaluation Engineer to help build and scale evaluation capabilities for our next generation of AI-powered this role you will develop evaluation frameworks datasets tooling and quality signals that enable teams to understand and continuously improve Generative AI and LLM-powered products. You will work closely with Machine Learning Software Engineering Quality Engineering Product Human Interface Data Science and domain experts to establish rigorous evaluation practices throughout the AI product will help define how we measure the quality of AI experiences across the Commerce domain including Store AI Shopping AI Learning AI Content GenAI Conversational AI and Platform is an opportunity to work at the intersection of machine learning software engineering data and product quality helping ensure our AI experiences are accurate relevant grounded reliable and useful for users around the world.

As a Machine Learning Evaluation Engineer you will design and build scalable evaluation systems for LLM Generative AI Conversational AI and Agentic AI will:nn Design and develop automated evaluation frameworks and pipelines for AI-powered Define evaluation methodologies and quality metrics across dimensions such as accuracy relevance groundedness completeness consistency instruction following and task Build and maintain high-quality evaluation datasets including golden datasets benchmark sets regression suites adversarial scenarios and production-derived test Develop Auto Eval capabilities that enable teams to rapidly evaluate models prompts retrieval systems agents and end-to-end AI Design and implement model-based evaluation approaches including LLM-as-a-Judge while developing appropriate calibration and validation Develop Human-in-the-Loop (HITL) evaluation approaches for complex or subjective quality dimensions where automated evaluation alone is Define evaluation rubrics annotation guidelines grading criteria and quality standards in partnership with product teams domain experts and annotation Build mechanisms to calibrate automated evaluators against human judgment and measure evaluator consistency and Evaluate end-to-end AI systems including retrieval context construction prompts model responses tool use APIs and downstream product Develop evaluation methodologies for multi-turn conversations personalization recommendations tool use reasoning and agentic task Perform detailed error analysis and failure-mode investigation to identify opportunities for model prompt retrieval dataset and product Build reusable evaluation infrastructure APIs dashboards and developer tooling that can scale across multiple AI products and Integrate evaluation into development and CI/CD workflows enabling automated regression detection quality gates and release-readiness Connect offline evaluation results with production signals to continuously improve evaluation coverage and product Partner closely with Machine Learning Software Engineering Product Quality Engineering Human Interface and Data Science teams throughout research development evaluation launch and continuous improvement.

Typically requires a minimum of 7 years of related experience in Machine Learning Engineering ML Evaluation Software Engineering Data Science Quality Engineering or a related technical programming skills in Python and experience developing production-quality software ML systems data pipelines or evaluation developing or evaluating LLMs Generative AI Conversational AI NLP recommendation systems or other machine-learning-driven designing automated ML evaluation frameworks metrics benchmarks datasets or experimentation of modern LLM application architectures including prompting embeddings retrieval-augmented generation (RAG) tool use and agentic with model-based evaluation techniques and an understanding of the strengths and limitations of approaches such as with Human-in-the-Loop evaluation annotation or data-quality understanding of statistical analysis experimentation sampling and measurement performing model error analysis failure analysis and root-cause to work effectively across Machine Learning Engineering Product Quality and Data written and verbal communication skills with the ability to translate complex technical findings into clear actionable degree in Computer Science Machine Learning Artificial Intelligence Data Science Statistics Electrical Engineering or a related technical field or equivalent industry experience.

Experience building evaluation infrastructure for production-scale LLM or Generative AI evaluating RAG conversational systems AI agents personalization recommendations or multimodal building golden datasets regression suites automated quality gates and continuous evaluation integrating ML evaluation into CI/CD and production release with prompt evaluation model comparison experiment tracking and AI evaluating multilingual AI experiences across languages locales and with responsible AI evaluation including robustness safety bias and adversarial developing internal ML platforms developer tooling or self-service evaluation capabilities used across multiple working with large-scale datasets and distributed ML or data-processing degree in Computer Science Machine Learning Artificial Intelligence Data Science Statistics Electrical Engineering or a related technical field or equivalent industry experience.

Required Experience:

Unclear Seniority


About Company

Company Logo

Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more

View Profile View Profile