Senior Machine Learning Engineer, Agent Eval Platform
Santa Clara County, CA - USA
Job Summary
The Role
Moveworks AI agents dont just generate text they act. They plan call tools and change real state in enterprise systems on behalf of 5.5 million employees. That makes the central problem of our team an unusually hard measurement problem: how do you score what an agent did across a multi-step trajectory through a world it changed precisely enough that the score can teach it to do better
That signal is what this role owns. Youll build the judgement layer of our agent evaluation platform: the rubrics the judges the calibration against human labels the methodology that makes a score mean something. And the payoff is larger than a report card a judge good enough to grade a trajectory is a judge good enough to train against. The same calibrated signal that explains why an agent failed becomes the reward signal that stops it failing.
This isnt a pretraining role and it isnt a testing role. Its applied ML at a point where the methodology genuinely isnt settled: LLMs judging LLMs is an open research problem and were working it against agents that take real irreversible actions in stateful multi-tenant enterprise environments.
What you get to do in this role:
Judge design and calibration
- A shared base judge with per-item rubrics expressed as configuration next to the dataset so eval authors express intent rather than forking a prompt per eval
- Splitting the problem correctly: deterministic validators for checkable world state (was the ticket created with the right item routed to the right approver) and an LLM judge for the parts that are genuinely fuzzy was the clarifying question appropriate was policy followed was the path efficient
- Scoring that reports its own confidence so uncertain judgements route to a human instead of quietly becoming training data
- A standing calibration loop against human-labeled trajectories run in partnership with our annotation team they own the human labeling you own the calibrated judge artifact. How consistently humans agree with each other sets the ceiling on how good any judge can be so raising that ceiling is part of the job
- Fine-tuning a small judge model where an off-the-shelf one isnt good enough
- Guarding against correlated blind spots: our user simulator and our judge are both LLMs and they can be wrong in the same direction
- Offlineonline divergence: when simulation and production disagree being the person who can say why and keeping the suite re-seeded from new production failures so it cant quietly overfit
Self-learning for the agent harness
This is where the pillar is headed and a large part of why the seat exists.
- A calibrated trajectory judge is functionally a reward model. Turning ours into a process reward model a dense step-level signal for what a good agent trajectory looks like is the unlock
- Using that signal to optimize the agent itself: prompts tool selection planner behavior retrieval routing tuned against simulation rather than against production traffic
- Building the substrate a future RL effort runs on: versioned scenarios a repeatable simulated world and a reward signal calibrated to human judgement
- Holding the guardrail that keeps this honest: step-level scores train and diagnose; end-state outcomes are what we hold the agent to. Scoring individual steps is powerful for attribution and as a training signal and dangerously brittle as a definition of success
Qualifications :
To be successful in this role you have:
- 5 years in applied ML data science or ML-adjacent engineering with a track record of work that shipped and got used
- Experience turning subjective human judgement into a measurement that holds up one that other people and ideally other models can act on. This is the core of the job
- Strong applied ML fundamentals and comfort treating LLMs as a component you evaluate prompt and fine-tune rather than one you pretrain
- Strong Python and the discipline to ship production-grade code rather than notebooks
- Ability to think and communicate clearly about complex problems a large part of this job is convincing engineers that a number means what you say it means and being right
- A high degree of ownership and a bias toward shipping at startup pace
- Comfort with ambiguity and the judgement to know when a measurement is good enough to act on
Experience in at least 3 of these:
- LLM-as-judge or automated evaluation design and calibrating it against human judgement
- Human annotation programs: rubric authoring label quality and annotator throughput as a real constraint
- Search ranking recsys or online experimentation evaluation golden-set staleness offline/online divergence side-by-side rater agreement. This is the closest existing analog to agentic eval and it transfers directly
- Fine-tuning and evaluating small models: SFT preference tuning distillation
- Reward modeling RLHF/RLAIF or process reward models
- Agent trajectory analysis and step-level fault attribution
- Prompt engineering as an engineering discipline versioned tested and measured not tuned by vibes
Additional Information :
Work Personas
We approach our distributed world of work with flexibility and trust. Work personas (flexible remote or required in office) are categories that are assigned to ServiceNow employees depending on the nature of their work and their assigned work location. Learn more here. To determine eligibility for a work persona ServiceNow may confirm the distance between your primary residence and the closest ServiceNow office using a third-party service.
Equal Opportunity Employer
ServiceNow is an equal opportunity employer. All qualified applicants will receive consideration for employment without regard to race color creed religion sex sexual orientation national origin or nationality ancestry age disability gender identity or expression marital status veteran status or any other category protected by addition all qualified applicants with arrest or conviction records will be considered for employment in accordance with legal requirements.
Accommodations
We strive to create an accessible and inclusive experience for all candidates. If you require a reasonable accommodation to complete any part of the application process or are unable to use this online application and need an alternative method to apply please contact for assistance.
Export Control Regulations
For positions requiring access to controlled technology subject to export control regulations including the U.S. Export Administration Regulations (EAR) ServiceNow may be required to obtain export control approval from government authorities for certain individuals. All employment is contingent upon ServiceNow obtaining any export license or other approval that may be required by relevant export control authorities.
From Fortune. 2025 Fortune Media IP Limited. All rights reserved. Used under license.
Remote Work :
No
Employment Type :
Full-time
About Company
Learn here. Grow here. Make a difference here. At ServiceNow, our cloud?based platform and solutions deliver digital workflows that create great experiences and unlock productivity for employees and enterprises. Were growing fast, innovating even faster, and making an impact on our c ... View more