Senior Data Scientist — AI Evaluation & Quality | Internal AI Agents (Remote in Europe)
Job Summary
- Own and extend our offline/online evaluation suites across 20 internal AI agent processesdatasets (capability regression) LLM-as-a-judge rubrics and deterministic checks.
- Establish pre-launch quality gates: enforce pass/fail thresholds in CI/CD pipelines before agent prompt context or tool changes hit production.
- Work directly with domain experts to label cases and resolve annotator disagreementfixing definition criteria rather than averaging disagreement away.
- Build test datasets derived from real user & operational traffic (tickets internal chats colleague queries) rather than synthetic edge cases.
- Harden statistical methodology: handle judge drift verbosity bias non-determinism and measure true metric shifts vs. noise.
- Translate quality numbers into operational decisions: run weekly syncs with process owners to define clear quality vs. cost/latency trade-offs.
- 5 years in Data Science / Product Analytics / Applied AI roles with sustained product-level metric ownership.
- Production LLM Experience: In the last 12 years you have built shipped or evaluated LLM-based systems (RAG multi-step tool use agents) as a core primary job responsibility.
- Autonomous Quality Ownership: Proven track record of owning evaluation methodology or analytics for an entire product or end-to-end process (what to build vs. what NOT to build).
- Fluent Python & SQL: Ability to write clean data pipelines evaluation harnesses and dbt transformation models directly.
- Statistical Rigor: Applied knowledge of sampling hypothesis testing variance analysis and confidence intervals on noisy metrics.
- AI-assisted coding (Claude Code Cursor or Codex) is your default daily authoring environment for Python SQL and evaluation scriptsnot something you occasionally experiment with.
- You can walk us through concrete work tasks from the last month where AI coding tools accelerated your engineering and data analysis.
- AI-assisted coding is our default authoring environment not a bonus
- Claude Code is our main tool youll reach for it for SQL Python analyses dashboards and internal scripts
- Were looking for analysts who are already curious and fluent with AI coding or genuinely excited to become fluent fast
- We care about what you ship and how clearly you think
- If this idea excites you rather than worries you youll feel at home here
Required Experience:
Senior IC
About Company
What You Will Get In Return Make a genuine impact on the product Join our upward trajectory, and grow with us. We provide the resources and opportunities for continuous personal and professional development, empowering you to make a genuine impact on our evolving product. Work in the ... View more