Lead Voice AI Architect
Posted:
28 July 2026 (30+ days ago)
Application Deadline:
25 October 2026
Vacancies:
1 Vacancy
Job Summary
Key Responsibilities
Own end-to-end architecture: ASR NLU/LLM dialogue manager TTS telephony
Make build-vs-buy calls on ASR/TTS vendors (Deepgram Azure Speech ElevenLabs or open-source Whisper/Coqui)
Own latency budget across the pipeline (target sub-1.5s round-trip for live calls)
Mentor Conversational AI and ML/Voice engineers; own technical hiring bar
Must-Have Technical Qualifications
Production experience with at least one ASR engine (Whisper Deepgram Google STT Azure Speech) at scale
Hands-on experience with LLM-based dialogue systems (function calling RAG or fine-tuning)
Strong grasp of telephony integration: SIP WebRTC Twilio/Exotel/Ozonetel or similar
Experience optimizing for latency and cost simultaneously in a live voice pipeline
Comfortable working across Hindi/Hinglish and English accent handling for Indian/UAE markets
Track record of hitting production accuracy benchmarks: 95% transcription (WER-based) accuracy and 90% intent classification accuracy at scale not just in a demo
Experience handling dialect variation and code-switching (e.g. Hinglish Gulf Arabic variants) in production ASR/NLU
Good-to-Have
Prior experience at Amazon Alexa Microsoft Cortana/Speech Google Assistant or a voice-AI startup
Exposure to Arabic ASR/TTS for UAE market
2 years hands-on with LLMs/Generative AI in a production (not POC) setting
Screening Questions
Describe a voice pipeline youve architected where was the latency bottleneck and how did you fix it
How would you handle Hinglish code-switching in ASR for an Indian BPO client
Whisper vs a commercial ASR API how do you decide and whats the cost/accuracy tradeoff
How do you evaluate hallucination risk in an LLM-driven dialogue manager for a live customer call
Walk us through how you measured and hit a 95% transcription accuracy target in a past role what was the failure mode you had to fix to get there
Own end-to-end architecture: ASR NLU/LLM dialogue manager TTS telephony
Make build-vs-buy calls on ASR/TTS vendors (Deepgram Azure Speech ElevenLabs or open-source Whisper/Coqui)
Own latency budget across the pipeline (target sub-1.5s round-trip for live calls)
Mentor Conversational AI and ML/Voice engineers; own technical hiring bar
Must-Have Technical Qualifications
Production experience with at least one ASR engine (Whisper Deepgram Google STT Azure Speech) at scale
Hands-on experience with LLM-based dialogue systems (function calling RAG or fine-tuning)
Strong grasp of telephony integration: SIP WebRTC Twilio/Exotel/Ozonetel or similar
Experience optimizing for latency and cost simultaneously in a live voice pipeline
Comfortable working across Hindi/Hinglish and English accent handling for Indian/UAE markets
Track record of hitting production accuracy benchmarks: 95% transcription (WER-based) accuracy and 90% intent classification accuracy at scale not just in a demo
Experience handling dialect variation and code-switching (e.g. Hinglish Gulf Arabic variants) in production ASR/NLU
Good-to-Have
Prior experience at Amazon Alexa Microsoft Cortana/Speech Google Assistant or a voice-AI startup
Exposure to Arabic ASR/TTS for UAE market
2 years hands-on with LLMs/Generative AI in a production (not POC) setting
Screening Questions
Describe a voice pipeline youve architected where was the latency bottleneck and how did you fix it
How would you handle Hinglish code-switching in ASR for an Indian BPO client
Whisper vs a commercial ASR API how do you decide and whats the cost/accuracy tradeoff
How do you evaluate hallucination risk in an LLM-driven dialogue manager for a live customer call
Walk us through how you measured and hit a 95% transcription accuracy target in a past role what was the failure mode you had to fix to get there
Required Skills:
SIPAIWEBRTC