Franklin AI News Brief

ElevenLabs launches Eleven v4 and Turbo, with expressive speech and qualified latency claims

Key Takeaways

  • Eleven v4 and Turbo add delivery controls, multilingual speech and voice-cloning improvements.
  • ElevenLabs reports about 150ms median time to first speech for Turbo, excluding network latency, and says both models are available across its creative, agent and API products.
  • ElevenLabs has launched Eleven v4 and Eleven v4 Turbo, pairing a new expressive text-to-speech model with a lower-latency variant for conversational agents.
  • Its September 28 announcement, updated October 1, credits Mati Staniszewski and Piotr Dabkowski and says both models are available through ElevenAgents, ElevenCreative and ElevenAPI.
  • The company emphasizes how a line is delivered: its tone, pacing and emotional context, as well as whether a voice remains recognizable across a longer production.

ElevenLabs has launched Eleven v4 and Eleven v4 Turbo, pairing a new expressive text-to-speech model with a lower-latency variant for conversational agents. Its September 28 announcement, updated October 1, credits Mati Staniszewski and Piotr Dabkowski and says both models are available through ElevenAgents, ElevenCreative and ElevenAPI.

The company emphasizes how a line is delivered: its tone, pacing and emotional context, as well as whether a voice remains recognizable across a longer production. The launch also includes multilingual improvements and changes to voice cloning. Its performance figures remain publisher-reported measurements rather than results from an independent Franklin AI listening test.

More control over delivery and dialogue

Users can describe a delivery in natural language or add inline directions for particular phrases. The announcement gives examples of tags for laughter, accents and sound effects. ElevenLabs says v4 follows these directions more accurately than earlier models and has improved support for International Phonetic Alphabet pronunciation instructions.

The company also describes better consistency between speakers and across regenerated lines. Its intended use cases include narrated material and scenes where one character responds to another. These claims concern generated speech and the continuity of voice identity; they do not establish that a complete conversation system will answer correctly or carry out a user's request.

ElevenLabs reports that v4 was preferred by about 75% of listeners in blind comparisons against selected competing models. The accompanying note says listeners heard the same line from v4 and a competitor, then judged expressiveness and naturalness, with ties counted as half. This is a preference result under the stated comparison procedure, not a percentage of all speech tasks the model completes successfully.

Inworld is one provider represented in the comparison through its TTS-2 model. Franklin AI's provider overview explains its developer speech workflow and delivery controls; it does not independently validate ElevenLabs' preference result.

Turbo latency figures exclude a full application round trip

The announcement describes a median inference latency of approximately 100 milliseconds and a separate median time to first speech of approximately 150 milliseconds for Turbo. These refer to different measurements and should not be used interchangeably.

The note accompanying the time-to-first-speech figure defines it as the time from a request to audible speech. ElevenLabs measured identical scripts with default settings in September 2026, removed measured network latency for all systems and used WebSocket streaming for v4 Turbo.

A deployed voice agent still has an interaction beyond that measurement. Its network and any other components need to be assessed in the actual application. The published 150-millisecond figure therefore does not establish the total delay a customer will experience when asking a question and receiving an answer.

ElevenLabs says it optimized Turbo alongside ElevenAgents. That explains the product's intended integration, but the announcement does not offer a universal end-to-end latency guarantee for every third-party deployment.

More languages and shorter cloning samples

Both models support more than 90 languages, according to ElevenLabs. The company says a voice recorded in one language can speak supported languages while retaining its identity and adopting a native accent. These are broad product claims; the announcement's demonstrations do not establish equal quality for every language or voice.

Instant Voice Clones can use a 10-second recording, the company says. Eleven v4 also supports Professional Voice Clones, and ElevenLabs reports better similarity to the original voice and improved consistency across repeated generations.

For long-form work, the announcement describes more reliable request stitching, which chains generations into longer content. It connects that improvement with ElevenLabs Studio and the Reader App. A creator assessing the change should inspect both a single line and transitions between passages, since the claim covers continuity as well as the sound of one isolated output.

Availability is stated; detailed usage terms are not

ElevenLabs says both models are available now in its agent platform, creative products and API, and directs readers to create a free account to start generating. The announcement does not provide a complete price table, character allowance or per-plan cloning limits.

The strongest supported conclusion is that ElevenLabs is offering the models and describing specific improvements in control, expression and latency. Teams still need the relevant product terms and their own representative audio checks before estimating production cost or relying on a particular quality level. Neither the availability statement nor the preference score replaces that assessment.

Our read

Franklin AI Take

ElevenLabs removes network latency from its reported time-to-first-speech measurement. We would keep that qualification beside the headline latency figure and test the whole interaction before choosing a voice-agent stack. For creative work, repeated lines and passage transitions deserve as much attention as an impressive sample. The launch makes specific claims about both consistency and expression, so a useful evaluation should examine both rather than relying on one blind-preference percentage.