ElevenLabs Introduces Eleven v4 Text-to-Speech Model
A line of text can change significantly depending on how it’s spoken. “I need you to stay calm" should sound different depending on who's saying it, whether that's a doctor delivering it gently to a frightened patient, or a character in a game shouting to his squad before dropping into battle.
Today we're launching Eleven v4, our most emotive text-to-speech model yet, and its low-latency variant, Eleven v4 Turbo.
Ranked #1 by Artificial Analysis1, and preferred by ~75% of listeners in blind head-to-head tests over competing models2, Eleven v4 was designed to interpret tone, pacing, emotion, character, and context. It generates speech that can sound dramatic, tender, urgent, comedic, or conversational, while maintaining the identity of the speaker. More natural multi-speaker dynamics make conversations feel responsive, rather than like separate lines assembled together. Speakers respond to the context of the conversation, producing more natural dialogue and character interactions.
A new architecture built for performance
Built on an entirely new architecture, Eleven v4 is our most emotive Text to Speech model ranked #1 by Artificial Analysis1. Eleven v4 Turbo brings that same technology to low-latency use cases like agents. With a median inference latency of ~100ms, it can respond faster than the average pause between two people talking.
Both models can generate speech that feels emotive, rather than mechanical. Underlying audio fidelity is also higher across the board, with cleaner, more natural-sounding output.
Earlier generations of text-to-speech models can read text aloud, but Eleven v4 provides an emotional depth to speech that feels far more natural. The model can interpret the intended tone, how the speech should be paced, the character of the writing, and context from the text to deliver emotionally resonant speech.
Users also have fine-grained control over outputs, and can describe how a line should be delivered in natural language. They can also add specific instructions on how to say certain phrases, what emotion they should convey, and even sound effects, using inline tags like [laughs], [said angrily in French accent], [light rain], or [phone buzzing]. Eleven v4 follows these audio tags and direction prompts more accurately than prior models, so the delivery you describe is what you get. This makes it easier to direct narration, flesh out characters or their dialogue, and other places where delivery matters.
The ElevenLabs developer documentation covers the full tag syntax and how to apply these tags through the API. Support for the International Phonetic Alphabet (IPA) phonemes has also been significantly improved, so custom pronunciations behave more reliably.
Eleven v4 also brings improvements in consistency and dynamic conversations between multiple speakers. Using a new method for capturing speakers’ identities, Eleven v4 preserves the unique qualities of each voice. It’s able to keep speech consistent through agent conversations, audiobooks, or ads. Because the model understands the context of a whole scene, it generates natural dialogue where speakers respond to what's just been said, rather than stitching together isolated lines.
A Turbo model for low-latency use cases
High-quality voice models have tended to be slower to generate speech, meaning users have had to choose between fast or expressive voice agents. Many have opted for proficient, but monotone, agents that can seem robotic to a flustered customer just looking to resolve their issue.
Eleven v4 Turbo combines speed and emotion with a median time to first speech of ~150ms3, making it possible to deploy powerful agents in countless industries, from warm, reassuring agents that can accurately pronounce medical terms in healthcare, to fast-talking, slang-wielding agents to entertain players in gaming.