SWE (RL Environments) "Reinforcement Learning"
Job Summary
This dynamic company is helping push the frontier of LLMs and AI Agents through novel datasets and experimentation. We work on building the most complex infrastructure that powers frontier data creation for agentic and hard reasoning workflows. We work with all 5 of the leading AI labs and are becoming the go-to partner for data infrastructure for YC companies. We have a sharp hockey stick growth rate and are extremely talent-dense with most of our founding team coming from top IB and quant firms.
- Looking for recent graduates from top schools that are focused on excellence and depth rather than a lengthy track record.
- First author publications at top venues like NeurIPS ICML are highly desirable.
- Open to profiles from data companies with benchmarking experience.
- Ideal candidates should be able to create iconic benchmarks and have experience in supervised fine-tuning (SFT) or reinforcement learning (RL).
- Competitors are not publicly known to be in the same space; discretion is emphasized.
- Environment work is not publicly disclosed focusing on creating detailed simulations.
- Base salary ranges from $150k to $250k with significant bonus potential based on performance.
- Bonuses are uncapped and can significantly increase total compensation.
- Interview process includes a screener call take-home task and a two-day in-person work trial with decisions made within a week.
- Need candidates who are excited to work in a startup and take ownership of their work.
- Finding candidates who can quickly adapt and lead projects independently.
- Looking for candidates with a strong backend development focus and fast coding abilities.
- Preference for candidates from fast-growing startups or with experience in creating complex simulations.
- Flexible working hours with most team members in the office from noon to midnight.
- Work environment encourages getting the job done over strict hours.
- Visa sponsorships are possible with experience in H1B and other types
1 - 6 years of experience as a software engineer
MUST have explicit experience building reinforcement learning environments
One or more signals of excellence in software engineering:
At an RL company currently (see list)
Worked at a top VC backed startup
Great side projects
Quant hedge fund or similar
Previous experience as a founder or early-stage startup engineer
Track record of side projects published papers AI space
CS degree from top-30 school (US/CA/EUR only)
Strong fullstack engineer (depth in Python Typescript and other fullstack/backend languages)
Developed quantitative frameworks for measuring dataset quality/diversity.
Bias for action and execution; willing to do difficult and tedious work
Required Skills:
Python Agentic AI LLM AWS GCP
Required Education:
BS in engineering (EE preferred any engineering accepted