Opens in a new tab

ElevenLabs releases Eleven v4 and Eleven v4 Turbo

29-09-2026

ElevenLabs released Eleven v4 and Eleven v4 Turbo on 28 September 2026. Turbo starts speaking after about 150 milliseconds, against 262 for Cartesia Sonic 3.6 and 814 for OpenAI's GPT-4o mini TTS.

Written by:

Jorick van Weelie

Marketing Lead at DataNorth | AI Enthusiast & Tech Storyteller

elevenlabs releases eleven v4 and eleven v4 turbo with 150ms time to first speech
Sign up for our Newsletter

Published: 29 September 2026

ElevenLabs released Eleven v4 and Eleven v4 Turbo on 28 September 2026. The Turbo version starts speaking after about 150 milliseconds, against 262 milliseconds for Cartesia Sonic 3.6 and 814 milliseconds for OpenAI’s GPT-4o mini TTS on the same test. Both models speak more than 90 languages.

What can Eleven v4 and Eleven v4 Turbo do?

The two models split the job. Eleven v4 is built for audio you produce once and publish, such as an audiobook or a video voice-over. Eleven v4 Turbo is built for a voice that has to answer while someone is waiting, such as a phone agent or a live assistant.

The headline feature is control through audio tags. You write [laughs], [whispers] or [door slams] straight into the script and the model performs it. Eleven v3 had this, but v4 follows a sequence of tags more reliably, which matters when a single line needs two directions in a row.

Professional Voice Clones are back after being absent from v3. That is the version you train on a long recording of one speaker, and it is what studios and brands actually use. Instant clones still work from a 10-second sample.

What you need to know to try it:

  • More than 90 languages, with native accents rather than an English speaker reading foreign words
  • Up to 10,000 characters in a single generation
  • More than 17,500 existing voices work with v4
  • Output as MP3, WAV or PCM, and mu-law for telephony
  • Available in the web app, the API with REST and streaming, and ElevenAgents
  • Free tier of 10,000 credits a month, paid plans from $6 a month

Eleven v4 Turbo speed compared to Cartesia and OpenAI

Take this from the table: the gap to the nearest rival is not a few percent, it is a factor of nearly two, and against OpenAI it is more than five.

ModelMedian time to first speechWho makes it
Eleven v4 Turboabout 150 msElevenLabs
Cartesia Sonic 3.6262 msCartesia
OpenAI GPT-4o mini TTS814 msOpenAI

These are ElevenLabs’ own measurements, run in September 2026 over WebSockets with identical scripts, default settings and network time removed. That last part matters. Network time is what your users actually feel, and removing it flatters every model in the table, including the rivals. ElevenLabs separately reports a median inference latency of about 100 milliseconds, which is the model’s own compute time rather than the moment audio starts.

The company also says Eleven v4 ranked first on the Artificial Analysis Provider Voice Arena leaderboard for September 2026, and that roughly 75% of listeners preferred it in blind tests against five rival services. The leaderboard is third-party. The listener test is ElevenLabs’ own.

What ElevenLabs is not saying

There is no published price difference between v4 and v4 Turbo. The plans are quoted in credits per month, not in cost per minute of audio, so you cannot work out what a million characters of Turbo costs against a million characters of v4 without running it. For a call centre weighing a switch, that is the number that decides the business case.

Nor is there a quality comparison between the two. ElevenLabs says v4 is the quality model and Turbo is the fast one, but publishes no figure for how much expressiveness you give up by choosing speed.

What this means

Worth testing now if you run a voice agent that people talk to live. The concrete case is a small team running phone support or booking agents, where the pause before the bot answers is the single thing callers complain about. At about 150 milliseconds, Turbo is inside the gap people leave between turns in normal conversation, and the nearest alternative is not. Test it with your own network in the loop, because the published figure has network time taken out.

Worth watching rather than switching if you produce finished audio. The gains in v4 over v3 are real but incremental for pre-recorded work, where an extra 600 milliseconds costs you nothing. The one reason to move now is Professional Voice Clones, which v3 did not offer. If you park a brand voice on a cloned speaker, that alone justifies the upgrade, and the rest of the release is a bonus.

For more information, visit the official announcement of Eleven v4 on the ElevenLabs website.

Add DataNorth AI to your Google favorites