Enter a job title or keyword

AI Researcher ML Engineer (ASR & Speech Specialist)

LILT


Job Location:

Washington D.C., MD - USA

Monthly Salary: Not provided by the employer
Posted: 7 June 2026 (30+ days ago)
Application Deadline: 26 September 2026
Vacancies: 1 Vacancy
The job posting is outdated and position may be filled

Job Summary

About LILT

AI is changing how the world communicates and LILT is leading that transformation.

Were on a mission to make the worlds information accessible to everyone regardless of the language they speak. We use cutting-edge AI machine translation and human-in-the-loop expertise to translate content faster more accurately and more cost-effectively without compromising on brand voice or quality.

At LILT we empower our teammates with leading tools global collaboration and growth opportunities to do their best work. Our company virtuesWork together win together; Find a way or make one; Quicker than they expect; Quality is Job 1guide everything we do. We are trusted by Intel Corporation Canva the United States Department of Defense the United States Air Force ASICS and hundreds of global Enterprises. Backed by Sequoia Intel Capital and Redpoint were building a category-defining company in a $50B global translation market being redefined by AI.

Role Summary

We are seeking a highly skilled and visionary Senior AI Researcher / Machine Learning Engineer specializing in Automatic Speech Recognition (ASR) to anchor our core speech intelligence and benchmarking this role you will serve as our principal subject matter expert in AI speech data processing responsible for architecting training and scaling high-performance multilingual ASR models as well as developing rigorous quality benchmarks for agentic conversational AI.

A critical component of this position involves developing robust domain-adaptation frameworks that allow our models to dynamically incorporate proprietary customer terminology specialized industry jargon and multilingual nuances. You will collaborate with the Engineering Product and AI Research teams to transform state-of-the-art speech research into production-ready systems powering on-device real-time streaming translation and novel frontier model benchmarks.

Key Challenge: Scaling ASR models capable of dynamic vocabulary insertion for enterprise-grade ultra-low-latency real-time environments and end-to-end agentic AI benchmarking that goes beyond surface metrics.

This position may require access to U.S. government systems facilities or controlled information. Candidates must be U.S. citizens or nationals or otherwise eligible to obtain any required government access authorization. LILT will consider all applicants and will work with selected candidates to determine applicable requirements.

Key Responsibilities
  • Model Development & Innovation: Architect train fine-tune and evaluate state-of-the-art speech representations and ASR models (e.g. End-to-End Conformer Whisper RNN-T and hybrid CTC/Attention architectures) across multiple global languages.

  • Customization & Domain Adaptation: Design and deploy highly scalable algorithms for dynamic vocabulary insertion contextual biasing and language model (LM) personalization to precisely capture customer-specific terminology acronyms and product names.

  • Evaluation: Implement automated framework evaluations to benchmark model performance rigorously tracking Word Error Rate (WER) Character Error Rate (CER) embedding-based metrics latency budgets (RTF) and computing efficiency profiles under varying acoustic environments.

  • Agentic Benchmarking: Develop pioneering multilingual benchmarks for end-to-end conversational AI agents including speech-to-text and text-to-speech components and targeting the weaknesses of state-of-the-art frontier models.

  • Real-Time & Batch Speech Systems: Partner with core engineering teams to build optimize and maintain high-throughput pipelines optimized for both ultra-low latency real-time streaming inference and high-efficiency asynchronous (batch) multi-channel speech analysis.

  • Speech Pipeline Engineering: Develop and refine standard auxiliary components of the speech processing chain including Voice Activity Detection (VAD) speaker diarization punctuation restoration noise/acoustic normalization and audio pre-processing filters.

  • Cross-Functional Productization: Translate product requirements into technical AI roadmaps working hand-in-hand with Product Managers to ship speech-to-text simultaneous translation and semantic speech analytics features.


Required Technical Qualifications
  • Education: Masters or Ph.D. degree in Computer Science Electrical Engineering Computational Linguistics Data Science or a related quantitative field with an emphasis on speech processing or deep learning (or equivalent proven industry track record).

  • Speech Domain Expertise: Minimum of 35 years of dedicated professional experience developing ASR systems speech-to-text translation pipelines or advanced audio processing models.

  • Deep Learning Frameworks: Advanced proficiency with PyTorch or equivalent frameworks along with extensive experience utilizing dedicated speech toolkits such as Whisper NVIDIA NeMo Hugging Face Transformers Kaldi ESPnet or SpeechBrain.

  • On-device runtimes: Hands-on experience converting and running PyTorch models on at least one mobile inference runtime: ExecuTorch LiteRT (formerly TensorFlow Lite) or ONNX Runtime Mobile. You have personally taken a non-trivial model through conversion including resolving unsupported operations and dynamic-shape or decoder-loop issues.

  • Software & Infrastructure: Strong software engineering principles in Python with a clear understanding of data structures algorithm optimization and handling complex multilingual text/audio tokenization schemas.

  • Data Pipeline Mastery: Proven experience working with large-scale audio datasets audio augmentation techniques (e.g. SpecAugment noise injection) and text normalization/inverse text normalization (ITN) pipelines.

Preferred & Specialization Qualifications
  • High-Performance and on-device Inference: Experience optimizing models for constrained on-device and production environments using quantization (INT4/INT8/FP16) distillation ONNX Runtime TensorRT or Triton Inference Server.

  • Research Footprint: Peer-reviewed publications in premier speech and machine learning conferences (e.g. ICASSP INTERSPEECH NeurIPS ICLR ACL) are a strong plus or an active contribution footprint to open-source speech communities.

  • Hardware acceleration: Working knowledge of mobile NPU/DSP acceleration on the Android SoC landscape (Qualcomm QNN / Hexagon GPU and NNAPI delegates) and the trade-offs across Snapdragon MediaTek and Google Tensor.

  • Streaming Architectures: Deep technical familiarity with streaming neural architectures (e.g. block-processing streaming transformers or transducer models) and real-time network transport constraints (WebSockets gRPC).

  • Multilingual Engineering: Professional exposure to building zero-shot multilingual speech systems or managing cross-lingual acoustic phonology data.


Core Competencies & Soft Skills
  • Analytical Problem Solving: Ability to break down ambiguous business or product requirements into deterministic actionable machine learning experimentation frameworks.

  • Collaborative Communication: Strong capability to communicate intricate technical machine learning complexities to non-technical stakeholders across product design and executive leadership.

  • Ownership Mindset: Comfortable working in a fast-paced environment taking accountability from initial algorithmic hypothesis and exploratory research through to final production monitoring.

Our Story

Our founders Spence and John met at Google working on Google Translate. As researchers at Stanford and Berkeley they both worked on language technology to make information accessible to everyone. While together at Google they were amazed to learn that Google Translate wasnt used for enterprise products and services inside the quality just wasnt there. So they set out to build something better. LILT was born.

LILT has been a machine learning company since its founding in 2015. At the time machine translation didnt meet the quality standard for enterprise translations so LILT assembled a cutting-edge research team tasked with closing that gap. While meeting customer demand for translation services LILT has prioritized investments in Large Language Models human-in-the-loop systems and now agentic AI.

With AI innovation accelerating and enterprise demand growing the next phase of LILTs journey is just beginning.

Our Tech

What sets our platform apart:

  • Brand-aware AI that learns your voice tone and terminology to ensure every translation is accurate and consistent

  • Agentic AI workflows that automate the entire translation process from content ingestion to quality review to publishing

  • 100 native integrations with systems like Adobe Experience Manager Webflow Salesforce GitHub and Google Drive to simplify content translation

  • Human-in-the-loop reviews via our global network of professional linguists for high-impact content that requires expert review


LILT in the News

Information collected and processed as part of your application process including any job applications you choose to submit is subject to LILTs Privacy Policy at LILT we are committed to a fair inclusive and transparent hiring process. As part of our recruitment efforts we may use artificial intelligence (AI) and automated tools to assist in the evaluation of applications including résumé screening assessment scoring and interview analysis. These tools are designed to support human decision-making and help us identify qualified candidates efficiently and objectively. All final hiring decisions are made by people. If you have any concerns require accommodations or would like to opt-out of the use of AI in our hiring process please let us know at

LILT is an equal opportunity employer. We extend equal opportunity to all individuals without regard to an individuals race religion color national origin ancestry sex sexual orientation gender identity age physical or mental disability medical condition genetic characteristics veteran or marital status pregnancy or any other classification protected by applicable local state or federal laws. We are committed to the principles of fair employment and the elimination of all discriminatory practices.


Required Experience:

IC


About Company

Company Logo

Sequoia delivers the platform and guidance to help you hone a total people investment strategy that works well for your business and people.

View Profile View Profile