Enter a job title or keyword

Machine Learning Engineer (Real-Time Speech Translation)

LILT


Job Location:

Washington D.C., MD - USA

Yearly Salary: USD 120000 - 161434
Posted: 6 October 2026 (Yesterday)
Application Deadline: 3 January 2027
Vacancies: 1 Vacancy

Department:

Engineering

Job Summary

About LILT

AI is changing how the world communicates and LILT is leading that transformation.

Were on a mission to make the worlds information accessible to everyone regardless of the language they speak. We use cutting-edge AI machine translation and human-in-the-loop expertise to translate content faster more accurately and more cost-effectively without compromising on brand voice or quality.

At LILT we empower our teammates with leading tools global collaboration and growth opportunities to do their best work. Our company virtuesWork together win together; Find a way or make one; Dance in the customers shoes; Quicker than they expect; Quality is Job 1guide everything we do. We are trusted by Intel Corporation Canva the United States Department of Defense the United States Air Force ASICS and hundreds of global Enterprises. Backed by Sequoia Intel Capital and Redpoint were building a category-defining company in a $50B global translation market being redefined by AI.

Role Summary

We are building a new live translation product. We are looking for an ML Engineer to build the real-time speech translation backend that powers it.

You will own the real-time speech translation backend end-to-end from live audio input to translated output. You will build on LILTs production model serving platform (Ray Serve on GPU Kubernetes clusters) and our in-house adaptive machine translation models working closely with the senior architects of that platform and with our language processing researchers. The ASR and MT models exist. Your job is to make them work together as a low-latency streaming system that holds up in production.

This is a hands-on backend engineering role for those looking to own a real-time ML system from end to end supported by expert guidance and a clear product vision. Our team adopts an AI-first approach leveraging agentic coding and AI-driven PR reviews to accelerate development. We combine this with deep technical expertise requiring not only expert Python proficiency but also a comprehensive understanding of the entire ML stack from optimizing neural network architectures to managing production infrastructure on Kubernetes and making informed cost-aware decisions on hardware selection.

Location & eligibility: This position requires US citizenship and residence in the United States. Preferred locations are Washington D.C.; Boston MA; and Indianapolis IN (East Coast / ET timezone preferred).

Key Responsibilities
  • Real-time pipeline architecture: Build and manage services for high-throughput real-time audio and text streaming. Handle signal processing session lifecycles and concurrency management to ensure robust operation under load.

  • ML model integration: Integrate and serve streaming speech recognition and machine translation models collaborating with research teams to ensure models operate within required latency budgets.

  • Quality and confidence workflows: Develop logic for model-based confidence scoring routing segments for human intervention as needed and broadcasting real-time updates and corrections to end-users.

  • Infrastructure and scale: Architect and scale production ML infrastructure on GPU-accelerated Kubernetes clusters. Implement batching load balancing and autoscaling strategies to maintain performance and cost-efficiency.

  • Latency engineering: Establish comprehensive instrumentation for real-time performance. Identify bottlenecks optimize system throughput and drive down end-to-end latency metrics to meet production standards.

  • Interface and API definition: Define technical contracts and interfaces for audio ingestion and downstream service integrations. Partner with frontend and platform engineering teams to maintain clean robust integration points.

  • Collaboration and technical leadership: Drive cross-team alignment by defining clear API interfaces and technical contracts facilitating effective communication between engineering and product teams to ensure seamless system integration.

Required Qualifications
  • BS or MS in Computer Science or a related field or equivalent practical experience.

  • 3 years building production backend or ML serving systems in Python including strong async programming (asyncio) skills.

  • Hands-on experience with real-time streaming transport: WebSocket or gRPC bidirectional streaming session state backpressure and connection lifecycle handling.

  • Experience serving ML models in production on GPUs (Ray Serve Triton vLLM or similar) with Docker and Kubernetes.

  • Experience integrating speech or NLP models into production systems ideally streaming ASR (partial hypotheses endpointing VAD).

  • A latency-engineering mindset: you have profiled instrumented and optimized a real-time or low-latency system and can reason in per-stage budgets.

  • Effective use of AI coding agents (Claude Code Codex or similar) on top of fundamentals learned the hard way: you let agents do the typing but you can debug review and reason about every line without them and you know when not to trust them.

  • US citizenship and residence in the United States (contract requirement).

Preferred Qualifications
  • Ray Serve specifically including streaming responses and model multiplexing.

  • Familiarity with simultaneous or incremental MT concepts (retranslation prefix stability wait-k policies).

  • Machine translation quality estimation (COMET/CometKiwi class models) or other confidence estimation in production.

  • Message brokers for real-time fan-out and state distribution (RabbitMQ or similar).

  • Streaming text-to-speech integration and time-to-first-audio optimization.

  • WebRTC and SFU concepts or voice pipeline frameworks (LiveKit Agents Pipecat).

  • Handling of CJK and other non-Latin text in NLP pipelines (our first languages are Japanese Korean and English).

  • Observability tooling (Datadog Prometheus) for production ML systems.

Our Story

Our founders Spence and John met at Google working on Google Translate. As researchers at Stanford and Berkeley they both worked on language technology to make information accessible to everyone. While together at Google they were amazed to learn that Google Translate wasnt used for enterprise products and services inside the quality just wasnt there. So they set out to build something better. LILT was born.

LILT has been a machine learning company since its founding in 2015. At the time machine translation didnt meet the quality standard for enterprise translations so LILT assembled a cutting-edge research team tasked with closing that gap. While meeting customer demand for translation services LILT has prioritized investments in Large Language Models human-in-the-loop systems and now agentic AI.

With AI innovation accelerating and enterprise demand growing the next phase of LILTs journey is just beginning.

Our Tech

What sets our platform apart:

  • Brand-aware AI that learns your voice tone and terminology to ensure every translation is accurate and consistent

  • Agentic AI workflows that automate the entire translation process from content ingestion to quality review to publishing

  • 100 native integrations with systems like Adobe Experience Manager Webflow Salesforce GitHub and Google Drive to simplify content translation

  • Human-in-the-loop reviews via our global network of professional linguists for high-impact content that requires expert review


LILT in the News

Information collected and processed as part of your application process including any job applications you choose to submit is subject to LILTs Privacy Policy at LILT we are committed to a fair inclusive and transparent hiring process. As part of our recruitment efforts we may use artificial intelligence (AI) and automated tools to assist in the evaluation of applications including résumé screening assessment scoring and interview analysis. These tools are designed to support human decision-making and help us identify qualified candidates efficiently and objectively. All final hiring decisions are made by people. If you have any concerns require accommodations or would like to opt-out of the use of AI in our hiring process please let us know at

LILT is an equal opportunity employer. We extend equal opportunity to all individuals without regard to an individuals race religion color national origin ancestry sex sexual orientation gender identity age physical or mental disability medical condition genetic characteristics veteran or marital status pregnancy or any other classification protected by applicable local state or federal laws. We are committed to the principles of fair employment and the elimination of all discriminatory practices.


Required Experience:

IC


About Company

Company Logo

Sequoia delivers the platform and guidance to help you hone a total people investment strategy that works well for your business and people.

View Profile View Profile