AI Engineer, RL & Evals
San Francisco, CA - USA
Job Summary
San Francisco CA
On-site 5 days per week in-office.
Seed Stage / Well-Funded AI Startup
On-site 5 days per week in-office
$175000 $275000 Base
Flexibility to go higher for exceptional candidates.
Competitive Equity
Open to Visa Transfers including OPT and H-1B transfers. Other visa situations may be considered case-by-case.
27 years of experience as an AI Engineer Machine Learning Engineer Reinforcement Learning Engineer or strong backend/full-stack engineer with significant RL evals or ML post-training experience.
Full-time
2 candidates
This is a fast-growing seed-stage AI company building the data and infrastructure layer that helps improve AI model performance in subjective and difficult-to-evaluate domains.
The company works directly with leading AI research organizations and application-layer companies on post-training reinforcement learning environments evaluation systems and high-quality data infrastructure.
The engineering team operates at the intersection of backend engineering applied machine learning reinforcement learning and product development. Engineers have significant ownership over the systems and environments that help evaluate and improve modern AI models.
This is not a pure research role. The ideal candidate is a product-oriented AI engineer who enjoys building production systems shipping end-to-end software and working hands-on with RL environments evaluation frameworks agent systems and ML infrastructure.
You will work on technically challenging problems where traditional automated evaluation is difficult including creative and subjective domains where quality cannot always be measured through simple deterministic metrics.
The ideal candidate is someone who can move comfortably between backend engineering and applied ML take ownership of ambiguous problems and turn research concepts into reliable production systems.
- Build and scale reinforcement learning environments for complex and subjective domains.
- Design evaluation frameworks that measure model performance beyond traditional benchmark metrics.
- Create unique tasks grading systems and evaluation methodologies for difficult-to-verify domains.
- Build systems that allow AI models and agents to be tested systematically.
- Develop infrastructure for evaluating model behavior quality reliability and performance.
- Design automated evaluation workflows that reduce reliance on manual human assessment.
- Iterate on environments and evaluation systems based on model performance and research findings.
- Build reliable production infrastructure supporting RL and post-training workflows.
- Design and build production backend systems supporting AI and ML workflows.
- Build agent harnesses context layers APIs and supporting infrastructure.
- Develop data pipelines for scraping indexing embedding processing and evaluating large datasets.
- Build distributed systems capable of supporting large-scale ML and evaluation workloads.
- Develop infrastructure connecting models environments datasets agents and evaluation systems.
- Work across backend engineering and ML systems to turn ideas into production-ready products.
- Improve system scalability reliability performance and maintainability.
- Own backend and infrastructure components from architecture through production.
- Ship production code end-to-end rather than working exclusively on research prototypes.
- Build AI-powered products and infrastructure used by internal teams and external partners.
- Develop systems that improve the quality and usefulness of AI models in real-world applications.
- Build agentic systems and evaluation infrastructure for environments where traditional automated metrics are insufficient.
- Translate ambiguous product and research requirements into practical engineering solutions.
- Work across product engineering and research requirements to deliver production systems.
- Rapidly prototype validate and productionize new ideas.
- Balance technical experimentation with reliability and production quality.
- Collaborate with internal research teams on RL post-training evaluation and model improvement.
- Work directly with leading AI labs and technical partners to develop environments and evaluation frameworks.
- Translate research concepts into production engineering systems.
- Help define tasks environments evaluation methodologies and technical requirements.
- Communicate technical decisions and tradeoffs clearly across engineering and research teams.
- Contribute to technical strategy around AI evaluation and post-training infrastructure.
- Work across backend ML data and product teams to solve complex AI problems.
- Take significant ownership over systems that directly influence model performance.
- 2 years of professional experience as an AI Engineer ML Engineer RL Engineer or strong software engineer working on AI systems.
- Strong production engineering experience.
- Experience building reinforcement learning environments evaluation systems or ML post-training infrastructure.
- Experience shipping production software rather than working exclusively on research.
- Strong backend or full-stack engineering experience.
- Experience owning technical projects from design through production.
- Experience working with ambiguous technical problems.
- Experience working in fast-paced engineering-driven environments.
- Strong ability to bridge software engineering and applied ML.
- Experience collaborating with research product or engineering teams.
- Strong Python experience.
- Strong PyTorch experience.
- Hands-on reinforcement learning experience.
- Experience building ML evaluation frameworks or evaluation systems.
- Experience working with LLMs.
- Strong backend engineering fundamentals.
- Experience with distributed systems.
- Experience building data pipelines.
- Experience with APIs and production services.
- Experience with data processing indexing embeddings or retrieval systems.
- Strong testing and debugging practices.
- Experience building production ML or AI infrastructure.
- Experience building reinforcement learning environments.
- Experience designing or implementing evaluation frameworks.
- Understanding of RL concepts and practical application.
- Experience evaluating LLM or agent behavior.
- Experience building systems for model post-training.
- Experience designing tasks benchmarks or grading systems.
- Experience working with agentic systems or agent frameworks is a strong plus.
- Experience building agent harnesses or context layers is a strong plus.
- Comfortable working on domains where objective evaluation is difficult.
- Strong understanding of the interaction between models data environments and evaluation.
- Experience turning ML research concepts into production systems.
- Ability to work across model-facing and software infrastructure layers.
- High ownership and accountability.
- Strong technical judgment.
- Product-oriented mindset.
- Comfortable operating in ambiguity.
- Strong problem-solving and systems-thinking skills.
- Able to move quickly from prototype to production.
- Strong communication skills.
- Comfortable collaborating with research and engineering teams.
- Able to explain technical concepts clearly.
- Strong English communication skills.
- Curious and motivated by difficult AI problems.
- Comfortable working in a fast-paced startup environment.
- Willing to work on-site 5 days per week in San Francisco.
- $175000 $275000 base salary.
- Flexibility to go higher for exceptional candidates.
- Competitive equity.
- Full-time position.
- On-site work model 5 days per week in San Francisco.
- Open to visa transfers including OPT and H-1B transfers.
- Opportunity to work directly on RL AI evaluation post-training and AI infrastructure.
- Significant ownership over production AI systems.
- Opportunity to work with leading AI research organizations and technical partners.
- Opportunity to build infrastructure for difficult and emerging AI evaluation problems.
- Work at the intersection of AI engineering reinforcement learning evaluation and product development.
- Build the infrastructure that helps improve the performance of modern AI models.
- Work directly with leading AI labs and advanced AI teams.
- Solve difficult evaluation problems in subjective and creative domains.
- Own RL environments evaluation frameworks agent systems and backend infrastructure end-to-end.
- Work in a fast-growing seed-stage company with significant engineering ownership.
- Combine software engineering with hands-on applied ML work.
- Ship production systems rather than working exclusively on research prototypes.
- Work on technically challenging problems at the frontier of AI development.
- Help establish the technical foundation for a rapidly scaling AI company.
Required Experience:
IC
About Company
Senior software engineering jobs at top AI-native startups. Recruiting from Scratch advocates for candidates — 300+ placements, 29-day avg time to hire, 90+ NPS. Browse open roles.