Senior Staff ML Ops Engineer
Dallas, TX - USA
Job Summary
- Build and evolve our training infrastructure on Kubernetes with Infrastructure GPU scheduling autoscaling multi-node distributed jobs capacity strategy and the operators and workflow engines that keep long-running training reliable.
- Shape the developer-facing surface CLIs SDKs job submission templates paved paths designed with the teams wholl use them. Make the common case one command and keep the uncommon case possible.
- Shorten the inner loop. Time to first training run edit-to-signal latency local iteration before a job hits the cluster fast failure over slow mystery. Measure it publish it drive it down.
- Evangelize best-in-class tooling and frameworks. Track what the ecosystem is shipping evaluate honestly and make the case with working prototypes and migration paths or say plainly when a shiny thing isnt worth the switching cost.
- Strengthen the data and artifact layer. Dataset versioning sharding and high-throughput loading of large multimodal sensor data so jobs saturate GPUs instead of waiting on I/O.
- Turn one-off Python into durable tooling tested documented observable libraries CLIs and services with sane defaults and deletions where theyre overdue.
- Make experiments legible with the teams who live in them: experiment hygiene dashboards researchers trust a real model registry and lineage from dataset to checkpoint to simulation result.
- Ship CI/CD for models alongside autonomy and simulation so a model change is validated the same way a code change is.
- Build observability across the ML stack utilization throughput failure modes queue times cost per experiment. When a job fails at 3am on node 47 the researcher should find out why without you.
- Treat docs onboarding and support as product surface golden-path guides a new researcher productive on day twooffice hours that turn repeat questions into shipped fixes.
- Drive adoption not just availability. Prototype with real users watch them work iterate. A tool nobody adopts didnt ship.
- Make the platform boringly reliable fewer failures faster recovery and none of the manual steps that quietly cost a team days.
- Build guardrails that dont feel like walls with Security IT and Infrastructure: access controls data handling and cost governance that hold up in an IP-sensitive environment while staying self-serve.
- 5 years of software or infrastructure engineering including tools or platforms used by other engineers and operating ML or data-intensive production systems.
- Hands-on Kubernetes expertise GPU scheduling autoscaling Helm or equivalent networking fundamentals and the ability to debug a cluster under load rather than restart it.
- Excellent Python and a track record of designing APIs and CLIs other people enjoy using.
Practical AWS depth: object storage at scale IAM GPU compute networking cost management and infrastructure as code (Terraform Pulumi or similar). - Distributed training in PyTorch (DDP FSDP or similar) plus experiment tracking and model registry tooling from the perspective of someone who made them pleasant for others to use.
- Fluency with containers CI/CD and modern build systems including large monorepos.
- The ability to influence without authority: evaluate a framework on its merits pilot it credibly and persuade skeptical senior engineers to change how they work.
- A collaborative default youd rather co-own a system than draw a boundary around your part of it.
- User empathy: youd rather fix the third-most-interesting problem blocking ten people than the most interesting one blocking nobody.
- Strong product instincts strong writing and comfort operating autonomously in ambiguous territory.
- Passionate about self-driving technologies and frontier AI and about what a small world-class team can do with the right infrastructure.
- Internal developer platform research platform or DevEx work with a story about a tool whose adoption you grew from zero.
- Large-scale distributed GPU training: hundreds to thousands of accelerators NCCL high-performance cluster networking collective communication tuning.
- High-throughput loading of LiDAR or camera data and formats such as Parquet or WebDataset.
- Workflow and scheduling systems Argo Workflows Ray Flyte Kubeflow or Slurm.
- Build-system depth (Bazel or similar) including remote caching in a monorepo.
- Simulation infrastructure or large-scale batch evaluation pipelines.
- Background in ML robotics or autonomous systems infrastructure.
- Security- and IP-sensitive production environments.
- Open-source contributions to ML infrastructure or developer tools.
- Competitive compensation and equity awards.
- Health and Wellness benefits encompassing Medical Dental and Vision coverage (for full-time employees only).
- Unlimited Vacation.
- Flexible hours and Work from Home support.
- Daily drinks snacks and catered meals (when in office).
- Regularly scheduled team building activities and social events both on-site off-site & virtually.
- As we grow this list continues to evolve!
Required Experience:
Staff IC
About Company
The US yearly salary range for this role is: $184,000 - $272,000 USD in addition to competitive perks & benefits. Waabi (US) Inc.’s yearly salary ranges are determined based on several factors in accordance with the Company’s compensation practices. The salary base range is reflective ... View more