Member of Technical Staff, Inference Systems
Palo Alto, CA - USA
Job Summary
Company: Photon
Location: Palo Alto CA (on-site 5 days per week)
Compensation: $230000 - $350000 competitive equity
Employment Type: Full-time
Visa Sponsorship: Visa transfers and new sponsorship (H-1B TN)
Photon is building a next-generation AI inference platform from the ground up with a relentless focus on performance. It was founded by Stanford alumni with deep AI infrastructure experience including early work at Together AI and is engineering the entire inference stack with Rust at its core.
Photon has raised a $10M seed from notable investors and is currently in stealth ahead of announcing its fundraise and product.
Photon is hiring Members of Technical Staff (2 years) to build a high-performance inference platform from scratch. You are a systems engineer who knows inference internals (attention KV cache batching scheduling) and wants to own the whole stack rather than a narrow slice. You will join a small team building a new inference system in Rust where you shape every design decision.
- Build a new inference runtime in Rust owning batching scheduling request routing and the serving stack.
- Design KV cache management prefix caching and optimizations that reduce latency and cost per token.
- Scale serving across GPUs and nodes.
- Profile benchmark and ship performance improvements across the inference pipeline.
- Make core architecture decisions with the founding team.
- 2 years of systems engineering experience
- Deep knowledge of inference internals: attention KV cache batching and scheduling
- Experience inside inference engines such as vLLM SGLang or TensorRT-LLM
- Strong systems programming (Rust C or similar)
- Ability to work on-site in Palo Alto 5 days a week
Rust Python PyTorch C Go vLLM SGLang TensorRT-LLM CUDA Triton NCCL
REVENUE: 14% of first-year salary. Est. fee per hire $32K-$49K; 10 seat(s) up to $406K if all filled.
TARGET COMPANIES (suggested (vLLM/SGLang/TRT-LLM contributors)): Together AI Fireworks AI Baseten Modal NVIDIA Anyscale.
BEST-FIT CANDIDATE: 2 yrs; LLM inference internals; vLLM/SGLang/TRT-LLM; Rust/C systems; visa: transfers new H-1B/TN; location: Palo Alto 5 days.