Enter a job title or keyword

SWE (RL Environments) "Reinforcement Learning"

AI Talent Now


Job Location:

San Francisco, CA - USA

Monthly Salary: $ 300 - 400
Experience Required: 5years
Posted: 16 June 2026 (30+ days ago)
Application Deadline: 13 September 2026
Vacancies: 1 Vacancy

Job Summary

About us

This dynamic company is helping push the frontier of LLMs and AI Agents through novel datasets and experimentation. We work on building the most complex infrastructure that powers frontier data creation for agentic and hard reasoning workflows. We work with all 5 of the leading AI labs and are becoming the go-to partner for data infrastructure for YC companies. We have a sharp hockey stick growth rate and are extremely talent-dense with most of our founding team coming from top IB and quant firms.


This company builds the training data and evaluation infrastructure that frontier AI labs use to make their models better. We work with the worlds leading labs to design high signal datasets and run rigorous evaluations that go beyond static benchmarks. We are a small early team (post Series A) where individual contributors have a direct impact on how the next generation of models learn and improve.

As an RL Environment Engineer youll design datasets that directly influence how frontier models learn and work hands-on with research teams at top AI labs.

  • Looking for recent graduates from top schools that are focused on excellence and depth rather than a lengthy track record.
  • First author publications at top venues like NeurIPS ICML are highly desirable.

Candidate Requirements

  • Open to profiles from data companies with benchmarking experience.
  • Ideal candidates should be able to create iconic benchmarks and have experience in supervised fine-tuning (SFT) or reinforcement learning (RL).

Company Overview

  • Competitors are not publicly known to be in the same space; discretion is emphasized.
  • Environment work is not publicly disclosed focusing on creating detailed simulations.

Compensation and Logistics

  • Base salary ranges from $150k to $250k with significant bonus potential based on performance.
  • Bonuses are uncapped and can significantly increase total compensation.

Timeline and Urgency

  • Interview process includes a screener call take-home task and a two-day in-person work trial with decisions made within a week.

Pain Points

  • Need candidates who are excited to work in a startup and take ownership of their work.
  • Finding candidates who can quickly adapt and lead projects independently.

Ideal Candidate Profile

  • Looking for candidates with a strong backend development focus and fast coding abilities.
  • Preference for candidates from fast-growing startups or with experience in creating complex simulations.

Work Arrangement

  • Flexible working hours with most team members in the office from noon to midnight.
  • Work environment encourages getting the job done over strict hours.

Visa Sponsorship

  • Visa sponsorships are possible with experience in H1B and other types


Requirements

1 - 6 years of experience as a software engineer


Work experience

MUST have explicit experience building reinforcement learning environments

Updated

One or more signals of excellence in software engineering:

  • At an RL company currently (see list)

  • Worked at a top VC backed startup

  • Great side projects

  • Quant hedge fund or similar


Previous experience as a founder or early-stage startup engineer


Track record of side projects published papers AI space

Education

CS degree from top-30 school (US/CA/EUR only)


Hard skills

Strong fullstack engineer (depth in Python Typescript and other fullstack/backend languages)


Developed quantitative frameworks for measuring dataset quality/diversity.


Soft skills

Bias for action and execution; willing to do difficult and tedious work




Required Skills:

Python Agentic AI LLM AWS GCP


Required Education:

BS in engineering (EE preferred any engineering accepted