Remote job
Software Engineer, Voice - Milan
Job details
About this role
Role overview A specialist engineering role owning the real-time voice layer of a conversational AI platform that already runs in production for enterprise clients. The mission is to push past "works reliably" into "feels human on a real phone line," with end-to-end ownership of voice architecture, model and provider choices, the latency budget, and the conversational feel. This is the first full-time engineer dedicated to voice, with the scope to grow into leading the function over time.
Responsibilities - Own the real-time voice pipeline end to end: telephony audio ingress, streaming STT, turn-taking, the agent brain, and streaming TTS output - Engineer perceived latency through semantic end-of-turn detection, preemptive generation on partial transcripts, eager TTS, and filler and backchannel cues that mask tool calls - Make turn-taking feel human via barge-in that survives noisy lines and endpointing policies tuned to dialog state, so callers are never cut off mid-utterance - Raise voice quality on real phone audio: benchmark and A/B STT and TTS providers on G.711 calls (Italian first), exploit wideband or HD voice where carriers allow it, and trial context-aware TTS and conversational speech models - Build a trustworthy evaluation harness with per-stage latency budgets, turn-taking metrics, regression suites on recorded calls, and quality gates before anything reaches clients - Keep production reliable with per-stage observability, live-call telemetry, and CI/CD discipline
Requirements - Hands-on background with telephony infrastructure: SIP trunking, SBCs, and enterprise telephony platforms - Practical ML audio skills, including evaluating or fine-tuning ASR and TTS models and working with speech datasets - Proficiency in Elixir, since the agent platform is built on it - Open-source contributions to voice or audio projects - A latency obsession: reasoning in milliseconds per stage, instrumenting before optimizing, and understanding measured versus perceived latency - A product ear: ability to translate what a conversation sounds like into engineering priorities and measurable evaluations - Daily fluency with AI-native coding tools (for example, agentic coding assistants) - Fluent written and spoken English
Nice to have - Italian fluency, since the voice market is Italian-first and pronunciation, prosody, and evaluations are tuned weekly for it - Experience in the contact-center or CCaaS ecosystem
Benefits and work setup - Fully remote position, with optional use of a co-working office in Milan - Competitive compensation in the 40–70k RAL range plus a performance-based bonus - Meal vouchers and welfare programs to support everyday life - Dedicated education budget for ongoing learning - Formal career development planning - Top-grade hardware, potentially including a MacBook Air and iPhone - Company retreats in various locations throughout the year