Enter a job title or keyword

Machine Learning Performance Engineer Offboard Training & Inference

Applied Intuition


Job Location:

Sunnyvale, CA - USA

Yearly Salary: USD 215000 - 285000
Posted: 21 August 2026 (2 days ago)
Application Deadline: 18 November 2026
Vacancies: 1 Vacancy

Job Summary

Applied Intuition Inc. is powering the future of physical AI. Founded in 2017 and now valued at $15 billion the Silicon Valley company is creating the digital infrastructure needed to bring intelligence to every moving machine on the planet. Applied Intuition services the automotive defense trucking construction mining and agriculture industries in three core areas: tools and infrastructure operating systems and autonomy. Eighteen of the top 20 global automakers as well as the United States military and its allies trust the companys solutions to deliver physical intelligence. Applied Intuition is headquartered in Sunnyvale California with offices in Washington D.C.; San Diego; Ft. Walton Beach Florida; Ann Arbor Michigan; London; Stuttgart; Munich; Stockholm; Bangalore; Seoul; and Tokyo. Learn more at .

We are an in-office company and our expectation is that full-time employees primarily work from their Applied Intuition office 5 days a week. However we also recognize the importance of flexibility and trust our employees to manage their schedules responsibly. This may include occasional remote work starting the day with morning meetings from home before heading to the office or leaving earlier when needed to accommodate family commitments. This in-office expectation does not apply to contractor positions

About the Role

We are looking for a performance engineer who specializes in making large-scale machine learning workloads fast and cost-efficient in the datacenter. This role is focused on distributed training runs spanning many nodes and high-throughput batch inference sweeping petabytes of real-world autonomy logs for auto-labeling data mining ground-truth generation and evaluation.

The optimization target here is not tail latency on a vehicle - it is throughput cluster goodput and cost per unit of data processed. A training run that wastes 30% of its GPU-hours on stalled data loaders or an offline inference sweep that takes a week instead of a day directly slows down how fast the whole company can iterate. You will own the gap between what our fleet of accelerators is theoretically capable of and what our workloads actually achieve: profiling across the stack finding where the compute and the wall-clock time actually go and closing the difference.

You will work at the intersection of accelerators ML frameworks and large-scale data infrastructure partnering with the teams who own each layer to land wins that show up in training time-to-result and offline processing cost. At Applied we encourage all engineers to take ownership over technical and product decisions closely interact with users to collect feedback and contribute to a thoughtful dynamic team culture.

At Applied you will:

  • Profile and optimize distributed training end to end - data loading and preprocessing augmentation kernel execution gradient communication and checkpointing

  • Optimize large-scale offline and batch inference over petabyte-scale sensor logs: batching and scheduling strategies quantization and low-precision execution graph optimization and accelerator saturation across long-running sweeps

  • Establish roofline and performance models for our workloads quantify the gap between achieved and theoretical performance and stack-rank optimization opportunities by impact and effort

  • Improve multi-node scaling efficiency: sharding and parallelism strategies collective communication interconnect utilization and memory-bandwidth and kernel-fusion bottlenecks

  • Drive cluster goodput - reduce GPU idle time from input pipeline stalls storage and network I/O scheduling gaps stragglers and failure recovery on long-running jobs

  • Build the benchmarking observability and regression-detection tooling that keeps performance from silently degrading as models and code evolve

  • Collaborate with engineers across functions to solve complex data and compute problems at scale

  • Contribute to a team culture that values effective collaboration technical excellence and innovation

Were looking for someone who has:
  • Hands-on ML performance engineering experience: profiling roofline analysis throughput optimization and root-cause investigation in production systems

  • Experience with distributed multi-node training at scale (FSDP DeepSpeed Megatron NCCL or equivalent) including diagnosing scaling inefficiency as node count grows

  • Deep familiarity with GPU or accelerator performance concepts - memory bandwidth kernel launch overhead occupancy quantization collective communication

  • Experience with high-throughput or batch inference systems (NVIDIA Triton Inference Server TensorRT ONNX Runtime Ray or similar)

  • Fluency in Python and proficiency in C or another systems language

  • Excellent debugging analytical and problem-solving skills

  • A deep understanding of machine learning foundations and the ability to develop technical solutions for problems with no established playbook

Nice to have:

  • GPU kernel development experience: CUDA Triton CUTLASS or hand-tuned attention implementations

  • Experience with profiling toolchains such as Nsight Systems/Compute PyTorch Profiler or perf

  • Experience with GPU scheduling and orchestration on Kubernetes Slurm or Ray including multi-tenant cluster utilization

  • Experience with fault tolerance and elastic training for long-running jobs - checkpointing strategy straggler mitigation preemption recovery

  • Familiarity with autonomy or robotics data (ROS OpenCV multi-sensor log formats)

Dont meet every single requirement If youre excited about this role but your past experience doesnt align perfectly with every qualification in the job description we encourage you to apply anyway. You may be just the right candidate for this or other roles.

Applied Intuition is an equal opportunity employer and federal contractor or subcontractor. Consequently the parties agree that as applicable they will abide by the requirements of 41 CFR 60-1.4(a) 41 CFR 60-300.5(a) and 41 CFR 60-741.5(a) and that these laws are incorporated herein by reference. These regulations prohibit discrimination against qualified individuals based on their status as protected veterans or individuals with disabilities and prohibit discrimination against all individuals based on their race color religion sex sexual orientation gender identity or national origin. These regulations require that covered prime contractors and subcontractors take affirmative action to employ and advance in employment individuals without regard to race color religion sex sexual orientation gender identity national origin protected veteran status or disability. The parties also agree that as applicable they will abide by the requirements of Executive Order 13496 (29 CFR Part 471 Appendix A to Subpart A) relating to the notice of employee rights under federal labor laws.

FOR US-BASED ROLES: Applied Intuition is committed to providing an accessible and inclusive application and interview experience to applicants who are disabled veterans and other applicants with disabilities or medical conditions. Reasonable accommodations are available requesting an accommodation will not affect your candidacy in any way and you are not required to disclose the nature of your disability or medical condition in order to make a request.
If you require an accommodation please contact . We will work with you!


Required Experience:

IC