AI Inference Engineer
Palo Alto, CA - USA
Job Summary
About the Role
Premier Global Links LLC is seeking an experienced Member of Technical Staff Inference Systems to build and optimize a high-performance AI inference platform from the ground up.
This role is focused on LLM inference model serving distributed systems and inference runtime performance. The ideal candidate has hands-on experience with production inference systems and strong systems engineering skills with Rust experience highly valued.
Key Responsibilities
- Build and optimize production LLM inference and model-serving systems.
- Develop inference runtime components using Rust and other systems-level technologies.
- Design and implement batching scheduling request routing and serving infrastructure.
- Build and optimize KV cache and prefix caching systems.
- Scale inference workloads across multi-GPU and multi-node environments.
- Profile benchmark and optimize latency throughput reliability and cost.
- Work with inference engines such as vLLM SGLang or TensorRT-LLM.
- Investigate performance bottlenecks across the inference stack.
- Contribute to core architecture and technical decisions for the platform.
- Collaborate with a small hands-on engineering team in a fast-paced environment.
Required Qualifications
- 210 years of experience in backend distributed systems or systems engineering.
- Hands-on experience building operating or optimizing LLM inference or serving systems.
- Deep understanding of transformer inference internals including attention KV cache batching and scheduling.
- Experience with a production inference engine such as vLLM SGLang or TensorRT-LLM.
- Strong programming experience with Rust C Go or systems-level Python/PyTorch.
- Experience building performance-critical systems where latency throughput and cost are important.
- Strong distributed systems and production software engineering fundamentals.
- Ability to work on-site 5 days per week in Palo Alto CA.
Preferred Qualifications
- Production Rust experience.
- CUDA or Triton kernel development experience.
- Multi-GPU or multi-node serving experience.
- Experience with NCCL NVLink or RDMA.
- Experience with prefix caching speculative decoding or prefill/decode disaggregation.
- Contributions to open-source inference projects such as vLLM SGLang or Dynamo.
- Experience working on inference systems at an AI provider accelerator company research lab or similar organization.
- Bachelors or Masters degree in Computer Science Computer Engineering or a related technical field.
Technology Environment
Rust C Go Python PyTorch vLLM SGLang TensorRT-LLM CUDA Triton NCCL NVLink RDMA LLM Inference Distributed Systems
Compensation & Benefits
- $230000$350000 annual salary based on experience and qualifications.
- Equity opportunity starting at approximately 0.5% with flexibility based on experience.
- Professional growth and development opportunities.
- High-impact work within a fast-paced AI technology environment.
Work Arrangement
On-Site Palo Alto CA
Employees are expected to work from the Palo Alto office 5 days per week.
Equal Opportunity Employer
Premier Global Links LLC is an equal opportunity employer. Qualified applicants are considered based on their skills experience education and qualifications.