Enter a job title or keyword

Data Scientist

Kaizo


Job Location:

Amsterdam - Netherlands

Monthly Salary: Not provided by the employer
Posted: 25 September 2026 (3 days ago)
Application Deadline: 23 December 2026
Vacancies: 1 Vacancy

Job Summary

Job description

Do you want to work out how you actually measure whether an AI system is doing a good job and then make it better Were looking for a sharp analytical data scientist to own the evaluation and quality loop of our AutoQA product. You can be early in your career (recent graduates with strong LLM fundamentals are welcome) we have room and a clear path for you to grow.

In a nutshell
  • Join a fast-growing SaaS company in an international environment (steep learning curve guaranteed).

  • Own a high-impact problem: making LLM-powered quality assurance measurably accurate at scale.

  • Sit at the intersection of AI and CX working directly with enterprise customers and their real-world QA rubrics.

  • Grow into a Senior Data Scientist AI Engineer or ML Engineer role. We invest in progression.

  • Enjoy the perks: flexible hours open holiday policy an office in the heart of Amsterdam with hybrid flexibility visa sponsorship great gear workations and team events.

About Kaizo

At Kaizo we build a performance development and quality platform for customer support teams. Our AutoQA product uses LLMs to review support conversations against each customers own quality rubric automatically and at scale. Behind it sits a microservices-based stream processing platform handling over 200 million events per day (Kafka Kubernetes on Google Cloud ElasticSearch MongoDB BigQuery) and an LLMOps stack built around LangSmith for experimentation prompt management and tracing.

The hard part isnt calling an LLM. Its knowing with evidence how well the system performs on every customers unique rubric and having a reliable repeatable way to improve it. Thats where you come in.

What youll focus on

  • Translate customer rubrics into AutoQA instructions. Work with real customer quality criteria and turn them into precise testable instructions that LLMs can score reliably.

  • Run experiments that move accuracy. Design and execute evaluation experiments on large representative datasets using LangSmith and BigQuery and track quality with our performance metrics.

  • Build the datasets that make evaluation possible. Curate raw production data into golden datasets with balanced coverage and generate synthetic data to cover the rare cases that matter most. QA is a discipline of rare occurrences: distributions are skewed positives are scarce and resourcefulness beats volume.

  • Build LLM-as-a-judge pipelines to assess system quality internally and make evaluation repeatable.

  • Make results actionable. Your experiments should end in a decision: change a prompt adjust which tools the system uses surface context the AI is missing or flag where new capabilities are needed. Youll work with the AI team to ship those decisions.

  • Evaluate across the full stack. Beyond scoring quality youll help validate retrieval (RAG/IR) tool calling and speech pipelines (transcription and diarization quality).

  • Join customer calls with the team to understand how QA leaders define quality and feed what you learn back into the product.

What youll grow into
  • Shaping how customers monitor quality themselves catch drift and keep their AutoQA setup improving over time.

  • Smarter categorization and routing of conversations to power analytics and get the right tickets to the right evaluation.

Job requirements
  • A solid understanding of how LLMs work and hands-on experience prompting them for accuracy (coursework thesis internships or side projects all count; production experience is a bonus).

  • Good applied statistics: experiment design classifier evaluation precision/recall trade-offs and working with heavily imbalanced data.

  • Strong Python skills and fluency with the standard data toolkit (Pandas NumPy Jupyter). SQL is a plus.

  • An analytical evidence-first mindset: youd rather measure than assume.

  • Product sense and empathy for end users. Youll be building for QA managers and support agents not just for benchmarks.

  • Excellent written and verbal communication. Youll present findings to the team and join customer conversations.

  • 0 to 2 years of industry experience. Recent graduates with strong relevant work are encouraged to apply.

  • A team player whos comfortable wearing multiple hats. Were an early-stage company and things move fast.

Bonus points for:

  • Experience with LangSmith or similar LLMOps/evaluation tooling

  • Google Cloud Platform BigQuery Docker or Kubernetes

  • Speech/audio processing or ASR evaluation

  • Fine-tuning or building synthetic datasets for LLMs

Who youll work with

Youll join our AI team (two data scientists an AI engineer and an ML engineer) and collaborate closely with our data engineers frontend engineers designer product manager and CX teams. Youll have mentorship from day one and real ownership fast.

Whats in it for you
  • An office in the heart of Amsterdam with the flexibility of hybrid working

  • Visa sponsorship available for eligible candidates

  • Great office gear: MacBook tools desk chair whatever you need

  • Flexible working schedule and an open holiday policy

  • Fun workations and team events

  • A clear growth path into Senior Data Scientist AI Engineer or ML Engineer roles

All done!

Your application has been successfully submitted!

Youve already applied for this job

Thank you for your interest - weve already received your application so this new submission cant be accepted. Your previous application is on file.

If you need assistance or believe this is an error please email us at


Required Experience:

IC


About Company

Company Logo

Join our AI powerhouse for QA, instant insights, impactful coaching, and gamified engagement.Exclusive for Zendesk & Salesforce users.

View Profile View Profile