RL Environments Engineer
San Francisco, CA - USA
Job Summary
San Francisco California Primarily On-site
We are seeking an Agent Evaluation Infrastructure Engineer to build the environments evaluation systems and supporting infrastructure used to train and assess long-horizon enterprise AI agents.
You will work on the engineering and research problems behind realistic agent environments post-training systems and reliable evaluation of complex multi-step workflows.
- Design evaluation environments for long-horizon enterprise agent workflows.
- Define tasks state tools graders and reward signals used to evaluate and improve agents.
- Build high-fidelity representations of complex enterprise software environments.
- Develop infrastructure for rollouts orchestration trajectory inspection and grader pipelines.
- Measure both correctness and efficiency across multi-step agent behavior.
- Investigate evaluation failures reward-quality issues and agent behavior.
- Build production-quality systems rather than notebook-only research prototypes.
- Hands-on experience with AI environments evaluations reinforcement learning infrastructure or related agent-training systems.
- Strong software engineering fundamentals.
- Demonstrated ability to build and ship technical infrastructure.
- Understanding of evaluation methodology reward design graders and agent trajectories.
- Ability to work across languages and technology stacks based on system requirements.
A PhD is not required. Strong engineering and shipped environment or evaluation systems are more important than academic credentials or publication history.
The opportunity is open to exceptional new graduates early-career engineers and experienced senior candidates. Selection is based primarily on engineering strength and relevant technical work.
The role is anchored in San Francisco with a strong preference for in-person collaboration. Limited flexibility may be considered case by case for exceptional candidates.