Member of Technical Staff, Infrastructure
San Francisco, CA - USA
Job Summary
This is a hands-on infrastructure engineering role at an early-stage robotics company sitting at the intersection of distributed systems ML infrastructure and deployed robot fleets. You will own critical systems end-to-end from petabyte-scale data pipelines to GPU cluster orchestration and your decisions will shape the engineering foundation while the codebase is still young enough for them to matter.
- Build and own distributed systems for training orchestration compute scheduling fault tolerance and large-scale model training networking.
- Design data pipelines that ingest petabytes of robot fleet telemetry and video corpora keeping GPUs fed without stalls.
- Architect a continuous-learning loop that moves production trajectories from deployed robots back into training reliably and autonomously.
- Go wherever the bottleneck is week to week whether that is a deployment path a core library or an observability gap.
- Build internal tooling that compounds over time and makes every engineer and researcher on the team faster.
- 3 years of experience in systems engineering ML infrastructure or distributed systems ideally at a well-regarded tech company or research lab.
- Strong systems engineering fundamentals: proven experience building or maintaining distributed systems data pipelines or GPU and compute infrastructure.
- Hands-on experience shipping software on physically deployed robots or autonomous vehicles in production.
- Proficiency in Python and C/C.
- Experience with petabyte-scale data systems or multi-node training orchestration is a strong plus.
- Background in ML infrastructure or large-scale training infrastructure projects is a strong plus.
- Must be available to work on-site six days per week in San Francisco.
Base salary: $200000 to $350000 USD annually. Visa sponsorship is not available for this role.
On-site in San Francisco California United States. This role requires in-person attendance six days per week.