ElevenLabs has launched Eleven v4 and Eleven v4 Turbo, a pair of text-to-speech models the company describes as its fastest and most emotive yet. Announced September 28, the release splits the lineup between produced audio and real-time conversation: v4 targets narration, dubbing and expressive dialogue, while v4 Turbo is built for live voice agents with a company-reported median inference latency of about 100 milliseconds.
The headline feature for creators is directability. Eleven v4 supports inline audio tags placed directly in a script, such as [laughs], [whispers] or [door slams], letting a writer shape emotion, pacing and reactions without leaving the text. The model reads with awareness of who is speaking and what came before, supports more than 90 languages, multiple speakers, sound effects and generations up to 10,000 characters. Context stitching is designed to keep delivery consistent across long-form projects, and speaker stability aims to prevent vocal drift when a line is regenerated.
The release also restores Professional Voice Clones, the company’s highest-fidelity cloning option, which were unavailable in v3. All 17,500-plus voices in the ElevenLabs library work with v4, though older Instant and Professional Voice Clones need retraining. Both models can clone a voice from about 10 seconds of audio, and the company says voices can speak any supported language while retaining the original identity. Notably, SSML break tags are disabled in v4, with natural-language audio tags serving as the control mechanism instead.
The models are available through ElevenLabs’ creator app, agent platform and developer API, including on a free tier. A two-week introductory API offer was priced at $22 per million characters. The launch was ranked number one by Artificial Analysis on its speech leaderboard.
For creators, the practical move is to test v4 on a real script before committing a series to it: audition the same paragraph with different inline directions, check pronunciation of names and brand terms, and confirm the commercial-use terms for your plan. Direction is now part of the writing, so the creators who learn the tag vocabulary early will get performances that sound intentional rather than default.
Join the conversation
Load Facebook comments to read and reply using your Facebook account.