Enter a job title or keyword

Applied Scientist Research Engineer — Speech AI


Job Location:

San Francisco, CA - USA

Monthly Salary: Not provided by the employer
Posted: 13 June 2026 (30+ days ago)
Application Deadline: 10 September 2026
Vacancies: 1 Vacancy

Job Summary

Applied Scientist / Research Engineer Speech AI

HIGHLIGHTS

Location:San Francisco CA OR REMOTE
Position Type:Direct Hire
Salary:Based on Experience
Residency Status:US Citizen or Green Card Holder Only

We are looking for a senior technical contributor to help develop the next generation of real-time speech and conversational AI systems. This person will work across applied research model development training infrastructure evaluation and production deployment.

The role is ideal for someone who has deep experience with modern machine learning for audio speech and language and who enjoys moving beyond prototypes into systems that must perform reliably in live environments. You will work closely with engineering and product teams to improve model quality speed reliability and scalability for speech-driven user experiences.

This is a hands-on position. You should be comfortable experimenting with new model architectures training and evaluating large models improving inference performance and translating research results into production-ready capabilities.

Responsibilities

Develop High-Quality Speech Generation Systems

  • Build and improve machine learning models for natural expressive speech generation.
  • Work on controllability speaker consistency pacing tone and conversational timing.
  • Explore model architectures that improve output quality while keeping latency low.
  • Improve performance for real-time use cases where responsiveness and reliability matter.
  • Partner with infrastructure and product teams to move successful approaches into production.

Improve Speech Understanding and Recognition

  • Train adapt and evaluate models that convert speech into accurate usable text.
  • Improve recognition quality across varied speakers accents noisy conditions phone-quality audio interruptions and mixed-language conversations.
  • Use large-scale audio data weak labels self-supervised methods and targeted fine-tuning strategies.
  • Improve downstream usefulness of transcripts for conversation analysis structured output and intent understanding.

Advance Audio Representation and Compression Methods

  • Research and implement model components that represent speech efficiently and preserve perceptual quality.
  • Explore learned audio representations that support generation recognition and efficient deployment.
  • Evaluate different approaches for balancing quality speed compute cost and scalability.
  • Build systems that can support both experimentation and production use.

Build Training and Evaluation Workflows

  • Create reliable pipelines for collecting cleaning processing and evaluating speech data.
  • Design evaluation methods that combine automated metrics model diagnostics and human quality review.
  • Support large-scale training jobs across modern accelerator infrastructure.
  • Improve throughput reproducibility monitoring and cost efficiency of model development workflows.

Run Rigorous Experiments

  • Design controlled experiments to test model data and training improvements.
  • Compare approaches using clear benchmarks and production-relevant quality measures.
  • Move quickly from hypothesis to implementation measurement and iteration.
  • Communicate results clearly to research engineering and product stakeholders.

What Were Looking For

Machine Learning Depth

  • Strong background in modern machine learning especially speech audio language generative modeling multimodal systems or large-scale model training.
  • Ability to implement new model ideas efficiently and evaluate them with technical rigor.
  • Strong understanding of current techniques used in speech and language systems.

Speech and Audio Experience

  • Practical experience building or improving systems for speech generation speech recognition audio modeling or related areas.
  • Experience working with large and varied audio datasets.
  • Strong judgment around speech quality naturalness latency robustness and user-facing model behavior.
  • Familiarity with real-world audio issues such as background noise channel quality interruptions speaker variation and conversational dynamics.

Production Awareness

  • Experience training deploying or serving large models on modern compute infrastructure.
  • Understanding of practical inference constraints including latency memory use throughput quantization and serving efficiency.
  • Comfort working with systems that need to operate reliably in live low-latency environments.

Experimental Discipline

  • Experience designing benchmarks ablation studies data experiments and quality evaluations.
  • Ability to use both offline metrics and live product signals to guide model decisions.
  • Strong technical judgment when deciding which ideas are worth scaling and which should be abandoned.

Ownership and Execution

  • Comfortable working in a fast-moving technical environment with ambiguous problems.
  • Able to own projects from early exploration through deployment.
  • Strong collaboration skills across research engineering infrastructure and product.
  • High standards for model quality reliability and operational performance.

Ideal Background

The strongest candidate will have experience building speech or audio AI systems that are both technically advanced and practical to deploy. You should be motivated by improving natural conversation real-time responsiveness model robustness and scalable production performance.

You may come from a research lab AI startup speech technology company communications platform infrastructure team or another environment where speech models were trained evaluated deployed and improved at scale.

Nice to Have

  • Experience with distributed training for large models.
  • Publications patents or open-source work in speech audio language or machine learning.
  • Experience with real-time communication streaming audio contact-center technology or other latency-sensitive speech products.
  • PhD in Machine Learning Computer Science Electrical Engineering Artificial Intelligence or a related field; equivalent research or industry impact is also acceptable.

Benefits

  • Competitive health dental and vision coverage.
  • Equity participation.
  • Access to modern compute tooling and research resources.
  • Opportunity to work on advanced speech AI systems used in production environments.

We are GTN The Go To Network