← BlogBlog · TTS

ElevenLabs audio tags: what v2 and v3 actually do with them

Same Russian sentence, same voice (Vasiliy). Only the model and the tags change. English meaning: "Hi! [laughs] Are you serious? [whispers] I'll tell you a secret."

Published By Stanislav Shupilkin3 min read

Listen first: one sentence, four recordings

Multilingual v2 · English tags4.8 s
Text
«Привет! [laughs] Ты серьёзно? [whispers] Я расскажу тебе секрет.»
What it said
«Привет, Левс. Ты серьезно? Висперс, я расскажу…» — "Hi, Levs. Are you serious? Vispers, I'll tell…"
v3 · English tags3.9 s
Text
«Привет! [laughs] Ты серьёзно? [whispers] Я расскажу тебе секрет.»
What it said
«Привет. Ты серьёзно? Я расскажу тебе секрет.» — no tag words
v3 · Russian tags5.6 s
Text
«Привет! [смеётся] Ты серьёзно? [шёпотом] Я расскажу тебе секрет.»
What it said
«— Привет. — Ты серьёзно? — Я расскажу тебе секрет.» — no tag words
v3 · no tags3.0 s
Text
«Привет! Ты серьёзно? Я расскажу тебе секрет.»
What it said
«Привет. Ты серьёзно? Я расскажу тебе секрет.»

Short answer:in our test, Multilingual v2 doesn't treat audio tags as tags at all — it reads [laughs] out loud as "Levs" and [whispers]as "Vispers". Eleven v3 never spoke the tags, English or Russian, but whether it actually laughed or whispered is something our transcripts can't tell you — you have to listen to the players above.

What happened to the tag

ModelTagsWhat happened to the tagCharacters billed
Multilingual v2Englishread out loud as a word64
v3Englishnot in the transcript64
v3Russiannot in the transcriptnot recorded
v3none—44

If your tags "don't work", check the model first

We wanted laughter and whispers in our Telegram bot, so first we checked whether the model listens to tags. Every recording went back through speech recognition: if a tag gets pronounced, it shows up as text. On v2 it did — "Levs" and "Vispers" right in the middle of the sentence.

On v3 the tag words were gone, for [laughs] and for the Russian [смеётся]alike. The tagged recordings were longer though: 3.9 and 5.6 seconds against 3.0 without tags. What filled that time — a laugh, a pause, something else — a transcript can't show, honestly.

The brackets are not free. The tagged sentence billed 64 characters on both v2 and v3, the plain one 44. So on v2 you pay to hear "Levs".

What ElevenLabs itself says: the best practices guide describes audio tags like [laughs] and [whispers] in the v4 section and says the same techniques, tags included, apply to v3. It also warns that a tag lands more readily when the voice already has that delivery in its training data, and recommends testing with your own voice. On the models page, audio tags aren't listed among Multilingual v2 features. ElevenLabs now calls v3 a previous-generation model and recommends v4 — we haven't tested v4.

So if you hear the tag spoken, it's not a typo in the tag, it's the model

Questions

Why are my ElevenLabs audio tags read out loud?

In our test that's what Multilingual v2 did: square brackets are plain text to it. The same text on v3 came out without the tag words.

Do ElevenLabs v3 audio tags work in other languages?

We only tried Russian: v3 didn't pronounce [смеётся] or [шёпотом]. Whether it acted them out, our transcripts can't separate — listen to the third recording.

Do tags work in @smolevich_voice_bot?

Not yet. The bot's studio voice uses Multilingual v2 for English and Russian text and sends your text as is, so [laughs] will be spoken as a word. Leave bracket tags out for now.

How we tested
  • Date: August 18, 2026.
  • Models: eleven_multilingual_v2 and eleven_v3, plain ElevenLabs text-to-speech API call, identical voice settings.
  • Voice: Vasiliy (1REYVgkHGlaFX4Rz9cPZ), a Russian voice.
  • Texts: «Привет! [laughs] Ты серьёзно? [whispers] Я расскажу тебе секрет.» on v2 and v3; «Привет! [смеётся] Ты серьёзно? [шёпотом] Я расскажу тебе секрет.»on v3; the same sentence without tags on v3. We didn't try Russian tags on v2, and we didn't try English text.
  • How we checked: speech-to-text on every recording, file length, characters billed as reported by ElevenLabs.
  • Caveat:4 recordings, one voice, one sentence, one run each. Enough for "this is what we got", not for "the model always does this".