Talking to most AI avatars feels like addressing three systems taking turns.
Nuance, a Seattle startup founded last year by three former Apple researchers, has collected a $50M Series A led by Lightspeed Venture Partners.
Accel, Nvidia’s NVentures, South Park Commons and Define Ventures participated. The round follows a $10M seed.
CEO and co-founder Fangchang Ma argues today’s voicebots and avatars are assembled from separate speech-to-text, language model and text-to-speech stages, which stacks latency and discards human cues such as reactions and microexpressions.
Nuance is training a single model instead, with audio and video in and audio and video out. Ma says the training data cannot be scraped from the web and is gathered in-house or through partners: two-sided audiovisual recordings with cleanly separated soundtracks.
The company is pre-product and plans a public research preview later this year. Business use cases include interview practice, sales calls, customer support and language learning.
Why it matters: if one end-to-end model beats a stitched pipeline on latency, the economics of every AI avatar and voice agent reset.