Cloning a convincing voice has stopped being the hard part. Nuance and speed are the new battleground, and that is the ground ElevenLabs chose for its latest generation.
Unveiled on Monday, the v4 and v4 Turbo pair uses a rebuilt architecture that sharpens expression control, trims latency for voice agents and stretches language coverage past 90, up from 70 in the previous release.
The new architecture clones a voice from as little as 10 seconds of audio. Over long passages the model holds a speaker’s identity better and reads context to shift expression, an expansion of the inline tags introduced with v3. Users can now stack those tags and the model follows the sequence.
Latency is the enterprise hook. v4 can start generating audio as soon as the language model behind it begins producing an answer, which keeps agent conversations fluid, and it handles escalations, holds and confrontations differently. The company reports its biggest quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese.
Voice is getting crowded
Competitors include Cartesia, Deepgram, Fish Audio and WellSaid Labs, while Google and OpenAI keep improving their own models. ElevenLabs raised $500M from Sequoia earlier this year at an $11B valuation, with reports of a follow-on near $22B, and annualized revenue has climbed from roughly $330M to more than $600M. Over 55% of the business now comes from large companies.
Co-founder and chief executive Mati Staniszewski says an IPO is a goal for the next few years, without committing to a date.