SWE Task Evaluator Fully Remote
San Francisco, CA - USA
Job Summary
About the job
Mercor connects elite creative and technical talent with leading AI research labs. Headquartered in San Francisco our investors include Benchmark General Catalyst Peter Thiel Adam DAngelo Larry Summers and Jack Dorsey.
Position: SWE-Bench Task Auditor
Type: Contract
Compensation: $70$90/hour
Location: Remote
Role Responsibilities
- Evaluate the quality correctness and reproducibility of software-engineering benchmark tasks.
- Assess repository-level tasks reference patches test harnesses and grading integrity.
- Provide clear rubric-based written feedback to improve AI model training.
- Audit reference patches test runners and Docker isolation to detect answer leakage and reward hacking.
- Work independently and asynchronously to meet deadlines and enhance AI model performance.
Qualifications
Must-Have
- 3 years professional software engineering experience.
- Real open-source contribution or maintainer experience (merged PRs committer/maintainer roles).
- Strong ability to audit reference patches test runners and Docker isolation.
- Fluency across common ecosystems (Python and at least one of Java / Go / TypeScript / C).
Preferred
- Familiarity with SWE-Bench (Verified) or similar repository benchmarks.
- Maintainer history on major Python OSS (Django Flask scikit-learn sympy pytest etc.).
- Prior code-review or task-grading experience.
Application Process (Takes 2030 mins to complete)
- Upload resume
- AI interview based on your resume
- Submit form
Resources & Support
- For details about the interview process and platform information please check:
- For any help or support reach out to:
PS: Our team reviews applications daily. Please complete your AI interview and application steps to be considered for this opportunity.