Data Scientist
Amsterdam - Netherlands
Job Summary
Do you want to work out how you actually measure whether an AI system is doing a good job and then make it better Were looking for a sharp analytical data scientist to own the evaluation and quality loop of our AutoQA product. You can be early in your career (recent graduates with strong LLM fundamentals are welcome) we have room and a clear path for you to grow.
Join a fast-growing SaaS company in an international environment (steep learning curve guaranteed).
Own a high-impact problem: making LLM-powered quality assurance measurably accurate at scale.
Sit at the intersection of AI and CX working directly with enterprise customers and their real-world QA rubrics.
Grow into a Senior Data Scientist AI Engineer or ML Engineer role. We invest in progression.
Enjoy the perks: flexible hours open holiday policy an office in the heart of Amsterdam with hybrid flexibility visa sponsorship great gear workations and team events.
At Kaizo we build a performance development and quality platform for customer support teams. Our AutoQA product uses LLMs to review support conversations against each customers own quality rubric automatically and at scale. Behind it sits a microservices-based stream processing platform handling over 200 million events per day (Kafka Kubernetes on Google Cloud ElasticSearch MongoDB BigQuery) and an LLMOps stack built around LangSmith for experimentation prompt management and tracing.
The hard part isnt calling an LLM. Its knowing with evidence how well the system performs on every customers unique rubric and having a reliable repeatable way to improve it. Thats where you come in.
Translate customer rubrics into AutoQA instructions. Work with real customer quality criteria and turn them into precise testable instructions that LLMs can score reliably.
Run experiments that move accuracy. Design and execute evaluation experiments on large representative datasets using LangSmith and BigQuery and track quality with our performance metrics.
Build the datasets that make evaluation possible. Curate raw production data into golden datasets with balanced coverage and generate synthetic data to cover the rare cases that matter most. QA is a discipline of rare occurrences: distributions are skewed positives are scarce and resourcefulness beats volume.
Build LLM-as-a-judge pipelines to assess system quality internally and make evaluation repeatable.
Make results actionable. Your experiments should end in a decision: change a prompt adjust which tools the system uses surface context the AI is missing or flag where new capabilities are needed. Youll work with the AI team to ship those decisions.
Evaluate across the full stack. Beyond scoring quality youll help validate retrieval (RAG/IR) tool calling and speech pipelines (transcription and diarization quality).
Join customer calls with the team to understand how QA leaders define quality and feed what you learn back into the product.
Shaping how customers monitor quality themselves catch drift and keep their AutoQA setup improving over time.
Smarter categorization and routing of conversations to power analytics and get the right tickets to the right evaluation.
A solid understanding of how LLMs work and hands-on experience prompting them for accuracy (coursework thesis internships or side projects all count; production experience is a bonus).
Good applied statistics: experiment design classifier evaluation precision/recall trade-offs and working with heavily imbalanced data.
Strong Python skills and fluency with the standard data toolkit (Pandas NumPy Jupyter). SQL is a plus.
An analytical evidence-first mindset: youd rather measure than assume.
Product sense and empathy for end users. Youll be building for QA managers and support agents not just for benchmarks.
Excellent written and verbal communication. Youll present findings to the team and join customer conversations.
0 to 2 years of industry experience. Recent graduates with strong relevant work are encouraged to apply.
A team player whos comfortable wearing multiple hats. Were an early-stage company and things move fast.
Bonus points for:
Experience with LangSmith or similar LLMOps/evaluation tooling
Google Cloud Platform BigQuery Docker or Kubernetes
Speech/audio processing or ASR evaluation
Fine-tuning or building synthetic datasets for LLMs
Youll join our AI team (two data scientists an AI engineer and an ML engineer) and collaborate closely with our data engineers frontend engineers designer product manager and CX teams. Youll have mentorship from day one and real ownership fast.
An office in the heart of Amsterdam with the flexibility of hybrid working
Visa sponsorship available for eligible candidates
Great office gear: MacBook tools desk chair whatever you need
Flexible working schedule and an open holiday policy
Fun workations and team events
A clear growth path into Senior Data Scientist AI Engineer or ML Engineer roles
Your application has been successfully submitted!
Required Experience:
IC
About Company
Join our AI powerhouse for QA, instant insights, impactful coaching, and gamified engagement.Exclusive for Zendesk & Salesforce users.