Enter a job title or keyword

Member of Technical Staff, Evals

Handshake


Job Location:

San Francisco, CA - USA

Yearly Salary: USD 200000 - 350000
Posted: 1 October 2026 (14 hours ago)
Application Deadline: 29 December 2026
Vacancies: 1 Vacancy

Job Summary

About Handshake

Handshakes mission is to organize expert human knowledge to advance the AI economy. Handshake AI works directly with frontier labs on their most consequential data evaluation and post-training challenges building the systems that turn expert human knowledge into the data and evaluations that make frontier models better.

You will work alongside engineers researchers operators and builders from organizations including Scale AI Meta Google Amazon xAI Notion and Palantirand help build the systems that make expert human knowledge useful for advancing AI.

The Role

We are hiring a Member of Technical Staff Evals to help define how frontier AI systems are measured understood and improved. This is a broad high-ownership role for researchers who build.

You will partner with AI researchers domain experts and customers to develop new benchmarks reward and verifier systems agent-evaluation methodologies and data-quality techniques. You will work on questions at the center of frontier AI progress: what should be measured how to design evaluations that reflect real capability how to create high-signal feedback and how to build the environments and data systems that make those answers actionable.

Early members of the team will have unusual influence over our technical direction operating culture and the open-source software benchmarks and research products we build. We care more about demonstrated research capability technical judgment and a builders mindset than a specific title degree or career path.

Location: San Francisco preferred; we are open to exceptional candidates in other locations.

What youll do
  • Design and build evaluation frameworks benchmarks and methodologies for frontier LLMs AI agents multimodal models and reinforcement-learning environments.

  • Develop reward models programmatic verifiers graders and other feedback systems that make model behavior measurable and improvable.

  • Research what makes evaluations representative difficult reliable and resistant to shortcutting or reward hacking.

  • Build systems for high-quality human data including expert task design annotation methodologies data-quality signals and data-attribution techniques.

  • Run fast rigorous iteration loops: prototype evaluate interpret results diagnose failure modes and turn learnings into the next benchmark or system.

  • Publicly contribute to the field through benchmarks open-source tools research and technical writing.

What were looking for
  • PhD in ML/AI computer science data science or related fields (or equivalent research experience in industry).

  • Publications at top AI/ML venues like NeurIPS ICML ICLR COLM.

  • Builders who enjoy tinkering with agents and shipping high-quality software benchmarks or datasets (e.g. a strong GitHub profile / OSS contributions or product portfolio).

  • Strong Python skills experience building scalable software working with agents.

  • Strong knowledge of frontier AI: benchmarks eval techniques agent harnesses post-training recipes data shapes.

  • Comfort operating in an ambiguous fast-moving environment with substantial ownership.

Why join
  • Work at the very frontier of AI with most major AI labs researching some of the most important problems in Data and Evaluations.

  • Publish results and work in public through open-source benchmarks and software papers and blogs.

  • Join a rapidly growing company whose data business grew from zero to nearly $1B run rate in a year.

  • Help build an early technical organization where your work shapes the roadmap standards and culture.

  • Attend (and publish at) conferences like NeurIPS ICML ICLR COLM.

Perks

Handshake delivers benefits that help you feel supportedand thrive at work and in life.

The below benefits are for full-time US employees.

Ownership: Equity in a fast-growing company

Financial Wellness: 401(k) match competitive compensation financial coaching

Family Support: Paid parental leave fertility benefits parental coaching

Wellbeing: Medical dental and vision mental health support $500 wellness stipend

Growth: $2000 learning stipend ongoing development

Office: Commuting support free lunch and gym in our SF office

Time Off: Flexible PTO 15 holidays 2 flex days

Connection: Team outings & referral bonuses


About Company

Company Logo

The better career platform for Gen Z changing how, where, and why the next generation of talent builds their career.

View Profile View Profile