Staff Applied Scientist, AI Quality & Meta Evaluation
Seattle, WA - USA
Job Summary
As a Staff Applied Scientist on the Human Centered AI team you will be the technical engine behind our Data Quality Validation framework. This is a high-impact individual contributor role for a scientist who wants to architect and build not just advise. You will own the data science methodology underpinning our data quality validation models design the statistical frameworks that govern judge reliability and work hands-on to close the loop between automated evaluation and human ground will be the person who answers the hardest question in our stack: Can we trust the evaluators that are evaluating our modelsn
Design develop and iterate on the reasoning agent that serves as our adjudicator auditing Production LLM Judge outputs for hallucination drift and systematic biasnDevelop the statistical and ML approaches that detect when Production LLM Judges diverge from ground truth including confidence calibration entropy-based uncertainty quantification and out-of-distribution detectionnDefine the algorithms that determine what gets routed for deeper review moving the team from random sampling to principled risk-stratified smart samplingnDesign the hierarchical weighting model and the confidence interval framework that replaces misleading point estimates with statistically rigorous rangesnEstablish the standards for how immutable ground truth sets are built versioned and validated including inter-annotator agreement protocols nPartner with Autograder Developers to validate new LLM Judge through our standard validation processes ensuring LLM Judges are rigorously validated before reaching productionnServe as the scientific authority on data quality evaluation methodology for partner teams across ASE translating complex statistical findings into clear decision-readiness signals for engineering and leadership stakeholders
Masters degree in Statistics Data Science Machine Learning Computer Science or a related quantitative fieldn8 years of hands-on experience in applied data science ML research or evaluation sciencenDeep expertise in uncertainty quantification and model calibration including entropy modeling and Bayesian approachesnDemonstrated experience building disagreement detection or anomaly detection models in production ML systemsnStrong command of statistical measurement frameworks inter-rater reliability correlation analysis and statistical process controlnProven experience designing or contributing to Human-in-the-Loop (HITL) or active learning pipelinesnProficiency in Python for statistical modeling ML experimentation and data pipeline developmentnExceptional ability to translate rigorous statistical methodology into clear actionable guidance for engineering and product partners
PhD in Statistics Computer Science Machine Learning or a related fieldnExperience specifically in LLM evaluation science including autograder validation judge-as-a-model frameworks or RLHF data qualitynHands-on experience with large-scale reasoning models (e.g. 70B parameter models) used in chain-of-thought evaluation or meta-reasoning contextsnExperience defining governance gates or certification pipelines for AI systems in a CI/CD contextnFamiliarity with out-of-distribution detection techniques for identifying input drift in live production systemsnTrack record of publishing or presenting evaluation methodology work internally or externally
Required Experience:
Staff IC
About Company
Ask Siri to name the most successful company in the world and it might respond: Apple. And it's not just out of familial pride. Apple consistently ranks highly in profit, revenue, market capitalization, and consumer cachet. In 2018, the company became the first reach a trillion dollar ... View more