Agent Quality Evals Engineer 1754
Job Summary
This is a remote position.
Required Skills:
Experience evaluating ML LLM or non-deterministic systems. Strong test and benchmark design capability. Comfort working with noisy metrics thresholds and probabilistic behavior. Good scripting and automation skills. AI-First Expectations Uses AI to generate candidate eval cases and failure hypotheses but never confuses generated tests with validated quality. Approaches AI quality as an operating system not a QA afterthought.