Remote job
Voice AI Engineer
Job details
About this role
Role overview Build the real-time voice backbone of a conversational AI product. The role centers on engineering a streaming speech-to-text → language model → text-to-speech pipeline that must respond inside a sub-second budget, with natural-feeling barge-in, turn-taking, and recovery behavior. Expect hands-on work with audio media, transport protocols, and latency tuning against real call recordings.
Responsibilities - Design and ship the real-time voice pipeline that overlaps STT, LLM, and TTS within tight latency targets. - Implement barge-in detection, turn-taking logic, and graceful recovery so interactions feel human. - Own WebRTC and telephony transport integrations, plus a rehearsal sandbox for safe iteration. - Profile and tune latency, prosody, and interruption handling using real call data. - Collaborate with product and research to land voice-quality improvements end to end.
Requirements - Strong Python skills and prior experience with real-time audio, media pipelines, or streaming systems. - Comfort with latency profiling and low-level performance debugging. - Working knowledge of STT, TTS, and LLM APIs and an understanding of their tradeoffs. - A bias toward ownership and shipping production-quality code.
Nice to have - WebRTC, SIP/telephony, or digital signal processing experience. - Prior work on voice agents or conversational AI systems.
Benefits and work setup - Austin-based or remote, full-time engineering role with email-based applications reviewed by an engineer.