Enter a job title or keyword

Member of Technical Staff Pre-Training Infrastructure


Job Location:

San Francisco, CA - USA

Monthly Salary: $ 200000 - 350000
Posted: 13 July 2026 (30+ days ago)
Application Deadline: 10 October 2026
Vacancies: 1 Vacancy

Job Summary

Who is Recruiting from Scratch:
Recruiting from Scratch is a specialized talent firm dedicated to helping companies build exceptional teams. We partner closely with our clients to deeply understand their needs then connect them with top-tier candidates who are not only highly skilled but also the right fit for the companys culture and vision. Our mission is simple: place the best people in the right roles to drive long-term success for both clients and candidates.
of Technical Staff Pre-Training Infrastructure

Location: San Francisco CA
Company Stage of Funding: Seed Stage ($23M Raised)
Office Type: Onsite (5 Days Per Week)
Salary: $200000$.250.40% Equity

Company Description

Were representing a well-funded robotics and AI startup building autonomous systems for industrial environments. The company is developing a vertically integrated robotics platform that combines advanced machine learning robotics infrastructure and large-scale model training to solve some of the most challenging problems in physical automation.

As one of the earliest members of the pre-training organization youll play a critical role in building the infrastructure that powers large-scale foundation model training. This team sits at the intersection of distributed systems machine learning infrastructure and hardware optimization enabling researchers to train and iterate on increasingly sophisticated multimodal AI systems.

What You Will Do
  • Design and maintain distributed training infrastructure for large-scale foundation model development.
  • Build efficient and reproducible multi-GPU and multi-node training workflows.
  • Develop high-performance data pipelines capable of handling multimodal datasets including video and large-scale structured data.
  • Optimize GPU utilization training throughput and hardware efficiency across large compute clusters.
  • Implement systems for checkpointing experiment tracking evaluation reproducibility and model comparison.
  • Build scalable data loading sharding and preprocessing infrastructure to support rapidly growing datasets.
  • Debug and resolve issues across model code infrastructure networking storage and hardware layers.
  • Partner closely with research teams to accelerate experimentation and improve model training velocity.
  • Establish reliable training baselines and infrastructure standards that support future model development.
  • Help define the companys long-term training infrastructure strategy as one of the earliest hires in the function.
Ideal Candidate Background
  • 15 years of experience building infrastructure for large-scale machine learning training.
  • Direct experience owning or operating pre-training infrastructure for foundation models.
  • Experience managing distributed training systems across multi-node environments and clusters of 100 GPUs.
  • Deep understanding of training bottlenecks related to compute memory networking storage and data loading.
  • Extensive experience with PyTorch and/or JAX in production or research environments.
  • Strong systems engineering skills spanning machine learning infrastructure distributed systems and hardware optimization.
  • Proven ability to troubleshoot issues across the full stack including model code data pipelines infrastructure and hardware.
  • Experience working in highly autonomous environments with significant ownership and responsibility.
Preferred
  • Experience building infrastructure for multimodal video robotics or foundation model training.
  • Background supporting large-scale pre-training post-training RLHF preference learning or synthetic data workflows.
  • Experience with autonomous systems robotics autonomous vehicles or large-scale perception systems.
  • Familiarity with data quality systems dataset auditing deduplication and evaluation contamination detection.
  • Strong academic background in Computer Science Machine Learning Robotics or a related field.
  • Publications at top-tier machine learning conferences such as NeurIPS ICML or ICLR.
  • Experience as an early startup employee or sole owner of critical machine learning infrastructure systems.
Compensation and Benefits
  • Base salary: $200000$350000.
  • Equity: 0.250.40%.
  • Visa sponsorship available for select visa categories.
  • Opportunity to join as one of the earliest members of a highly technical machine learning infrastructure team.
  • Direct ownership of foundational systems that influence research velocity and model performance.
  • Exposure to cutting-edge robotics multimodal AI and large-scale foundation model development.
  • High-impact role with significant autonomy and technical ownership.
  • Collaborative onsite culture focused on speed execution and ambitious technical goals.
  • Opportunity to work alongside engineers and researchers from leading AI infrastructure and robotics organizations.
Salary Range: $200000-$350000 base.

About Company

Company Logo

Senior software engineering jobs at top AI-native startups. Recruiting from Scratch advocates for candidates — 300+ placements, 29-day avg time to hire, 90+ NPS. Browse open roles.

View Profile View Profile