Senior Software Development Engineer in Test — LLM Evaluation & Automation, T3E
Cupertino, CA - USA
Job Summary
You will build and maintain model level component or end-to-end evaluation coverage for the generative features our team validates. Your job is to leverage LLM judge scoring output quality in automation ensuring reliable repeatable eval jobs that run that produce actionable kinds of problems you will work on include:nn* Image / visual generation: validating model output and its associated classification metadata and detecting quality or behavior regressions across model updates.n* Natural-language generation: evaluating whether generated artifacts and responses match user intent moving at-desk LLM judges into a scalable and repeatable automation environment.n* Correctness beyond string matching: replacing exact-match checks for open-ended or factual responses with an LLM-as-judge stage integrated into the pipeline.n* Generated insights and summaries: assessing whether model-generated content is sensible and good enough to surface to will decide when a component-level check (an API or CLI that exercises the model against its framework) is sufficient and when a full end-to-end user flow is required and you will build the tooling for both.
Design build and maintain LLM-as-a-judge evaluation harnesses and integrate them into existing and new CI/automation and curate eval sets and rubrics; partner with modeling teams whose own eval sets can run hundreds of examples judged by a separate eval jobs at scale triage results and distinguish real model regressions from rubric problems or infrastructure noise so the signal stays and build test coverage and tooling from low level component tests through to end-to-end tests that exercise models and frameworks powering Generative AI featuresnPackage eval tooling for reuse reusable libraries and jobs that other engineers on the team and partner teams can quality gates and reporting so model regressions are caught and communicated before human eval or population with data scientists and modeling engineers on approaches to spot regressions across large output sets or between model updates.
BS in Computer Science Mathematics or a related field (or equivalent practical experience) nThree years of relevant industry experience in test automation software development or related areas.
Strong practical knowledge of Python including data-pipeline fluency (JSON/YAML REST APIs).nHands-on experience with LLM-as-a-judge evaluation and rubric design or a strong demonstrated ability to ramp into it software engineering fundamentals able to define atomic composable components and build maintainable pipelines and tooling not just debugging and triage skills; able to separate genuine regressions from infrastructure or rubric knowledge of the software development lifecycle testing methodologies and QA written and verbal communication; able to document clearly and describe quality signal to modeling and leadership to lead work across varying priorities and partner multi-functionally with modeling framework and infrastructure building on-device tooling and device/model eval integrating with CI/CD and job orchestration systems and comfort deploying tooling as reusable with generative model behavior image generation NLP or LLM output curating and reasoning about large datasets; comfort manually inspecting data (Jupyter or similar) to build intuition and drive next of dataset bias and fairness considerations in with database/query tooling (e.g. SQL) and dashboards/visualization for reporting quality with Xcode is a bonus.
Required Experience:
Senior IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more