← BlogBlog · STT

Armenian speech to text: what transcription models return

Short answer: of the seven models I ran, only four write Armenian in Armenian script — both Whisper sizes, Deepgram Nova-3 with the language set to hy, and ElevenLabs Scribe — and Whisper and Nova-3 get at least 43% of words wrong even on clean audio. Ignore Scribe's 2%: the recordings were voiced by ElevenLabs itself, so Scribe is listening to its own company's voice.

Published By Stanislav Shupilkin5 min read

Example: one Armenian recording, clean and in noise

These recordings are synthetic: the text was voiced by ElevenLabs (voice Sarah, model eleven_v3), not by a person. Noise was added by a script. Model output is shown exactly as returned. The last column is the share of the 44 reference words the model got wrong.

Recording 1 — clean23 s
Reference
Գրադարանը բացվում է առավոտյան ժամը իննին և փակվում երեկոյան յոթին։ Երկուշաբթի օրը մենք վերադարձնում ենք տասնհինգ գիրք և վերցնում ենք չորս նորը։ Ընթերցասրահում նստած են քսաներկու ուսանողներ, որոնք պատրաստվում են քննություններին։ Դասախոսությունը սկսվում է ժամը տասնմեկին, տևում է հիսուն րոպե, իսկ ընդմիջումը՝ քառորդ ժամ։
Whisper large-v3-turbo (Groq)words wrong: 63.6%
գրաթրանը բացվում է առավոցյան ժամը 9-ին և պակվում երեկոյան 7-ին։ 2-շապցի օրը մենք վերադարցնում ենք 15 գիրկ և վերցնում ենք 4 նորը։ ընդերցասրահում նստած են 22 ու սանողներ, որոնք պատրաստվում են կննություններին։ Դա սախոսությունը նսկսվում է ժամը 11-ին, տեվում է 50-րոպ է, իսկ ընդմիջումը կարորդ ժամ։
Deepgram Nova-3, language hywords wrong: 47.7%
Գրադարանը բացվում է առավոտյան ժամը 9ին 2շապտի օրը մենք վերադարձնում ենք 15 գիրք եւ Ընթերցասրահում նստած ենք 22 ուսանողներ, Դասախոսությունը սկսվում է ժամը 11ին, տեւում է 50 րոպե, իսկ ընդմիջումը՝ քառորդ ժամ,

The same clip from two other services: Deepgram Nova-3 in multimode returned “Gret herene bathsvume ara vodtien jamme innin, yvpackvume yerekoyan yottin…” (111%), Fish Audio returned “گر تهران بادسومه، آرزو تیانجامه…” (Arabic script, 136%).

Translation of the reference: “The library opens at nine in the morning and closes at seven in the evening. On Monday we return fifteen books and take four new ones. Twenty-two students are sitting in the reading room, preparing for exams. The lecture starts at eleven, lasts fifty minutes, and the break is a quarter of an hour.”

Recording 2 — noise, speech 8 dB louder than noise28 s
Reference
same text
Whisper large-v3-turbo (Groq)words wrong: 54.5%
Վրաթրանը բացվում է առավոցյան ժամը 9-ին և պակվում երեկոյան 7-ին։ Վերկու շապթի օրը մենք վերադարցնում ենք 15 գիրկ և վերցնում ենք 4 նորը։ ընդերց ասրահում նստած են 22 ուսանողներ, որոնք պատրաստվում են կննություններին։ Դասախոսությունը նսկսվում է ժամը 11-ին, տեվում է 50 հոպե, իսկ ընդմիջումը կարորդ ժամ։
Deepgram Nova-3, language hywords wrong: 43.2%
Գրադարանը բացվում է առավոտյան ժամը 9ին 2շապտի օրը մենք վերադարձնում ենք 15 գիրք եւ վերցնում ենք 4 նոր Ընթերցասրահում նստած են 22 ուսանողներ, Դասախօսությունը սկսվում է ժամը 11ին, տեւում է 50 րոպե, իսկ ընդմիջումը՝ քառորդ ժամ։
Recording 3 — noise 3 dB louder than speech, plus a ringing tone28 s
Reference
same text
Whisper large-v3-turbo (Groq)words wrong: 77.3%
Վրաթերանը բատվում է առավոցյան ժամը ինին։ Ես հակվում երեկոյան դյոցն։ Երկուշավցի օրը մենք վերադարդնում ենք 15 դիտ և վերտնում ենք 4 նորը։ ընդհերծասրահում նստաբ ենց շանգերկում իսանողնեք, բորոնք պատրասվում ենք հնություններին։ Շասախոսիթինը նսկսվում է ժամը 11-ին։ տեվում են հիչ ունգոտ է, իսկ ընդնիչումը խարող սմ։
Deepgram Nova-3, language hywords wrong: 54.5%
Գրադարանը բացվում է առավոտյան ժամը 9ին եւ փակում երեկոյան գյուտն է։ Երկուշաբթի օրը մենք վերադարձնում ենք 15ից եւ վերցնում ենք կոշտն օրը։ Ընկերցասրահում նստած են ցանկեեւ ուսանողները, Դասապոչությունը նա սկսվում է ժամը 11ին, տեւում է 50 րոպե, իսկ հմտությունը՝ քառորդությամբ։

Synthetic recordings: voice Sarah, ElevenLabs eleven_v3. Noise added by a script; the audio and every raw model response are in the benchmark repository.

Words wrong by model and noise level

Share of words wrong, median of three runs
Deepgram Nova-3, hy
Clean47.7%
+8 dB43.2%
−3 dB54.5%
Whisper large-v3 (Groq)
Clean59.1%
+8 dB95.5%
−3 dB75.0%
Whisper large-v3-turbo (Groq)
Clean63.6%
+8 dB54.5%
−3 dB77.3%
Deepgram Nova-3, multi
Clean111.4%
+8 dB127.3%
−3 dB100.0%
Fish Audio transcribe-1
Clean136.4%
+8 dB136.4%
−3 dB118.2%
Deepgram Nova-2
no Armenian
ElevenLabs Scribe v1 ⚠hears its own voice — not comparable
Clean2.3%
+8 dB2.3%
−3 dB2.3%
Share of words wrong, %

Two ways to get Armenian wrong

“Words wrong” here is word error rate (WER): the words a model substituted, dropped or added, divided by the words in the reference. That is why it can go above 100%.

The two usable engines fail differently. Whisper large-v3-turbo keeps every word but misspells a lot of them — “գրաթրանը” for “գրադարանը” (library); on the clean clip about one letter in five is off. Deepgram Nova-3 spells better but quietly skips parts of sentences: 12 of 44 words simply vanished from the clean transcript. Both write numbers as digits (“9-ին” for “իննին”) — counted as errors, though the meaning is right.

Moderate noise (+8 dB) barely moves turbo or Nova-3. At −3 dB Whisper starts inventing words, and Nova-3 swaps meanings: “ընդմիջումը” (break) became “հմտությունը” (skill). The full whisper-large-v3 is no safe upgrade: on the +8 dB clip it returned only the first sentence, three runs out of three.

The rest never got to Armenian: Nova-2 doesn't support Armenian — the API rejects language=hy(HTTP 400, “no such model/language/tier combination, try Nova-3”); only Nova-3 has Armenian in Deepgram's language list. Nova-3 multi writes Latin, Fish Audio writes Arabic script even with hy passed.

ModelClean+8 dB−3 dBOutput script
Deepgram Nova-3, hy47.7%43.2%54.5%Armenian
Whisper large-v3 (Groq)59.1%95.5%75.0%Armenian
Whisper large-v3-turbo (Groq)63.6%54.5%77.3%Armenian
Deepgram Nova-3, multi111.4%127.3%100.0%Latin
Fish Audio transcribe-1136.4%136.4%118.2%Arabic script
Deepgram Nova-2no Armenianno Armenianno Armenian—
ElevenLabs Scribe v1 ⚠2.3%2.3%2.3%Armenian

In plain terms: for Armenian, “does it support the language at all” comes before “how accurate is it”. Don't rank the Scribe row against the others.

Transcribing Armenian voice messages in the bot

@smolevich_voice_bot has two regular modes. “🚀 Fast” is the default: Whisper large-v3-turbo on Groq, 5 credits a minute. Fine for getting the gist; expect to fix the text. “🎯 Accurate” is ElevenLabs Scribe v1, 25 credits a minute — switch via “🎤 Transcription” → the current-mode button → “🎯 Accurate”. I have no fair number for it on real Armenian speech yet, so test it on your own clip — the free 200 monthly credits cover about 8 minutes. Skip “🪄 Exclusive” for Armenian — that one is Fish Audio.

Questions

Does Whisper support Armenian?

Yes, Armenian (hy) is in Whisper's language list. In this test the turbo model still got 55–77% of words wrong.

Which Deepgram model handles Armenian?

Per Deepgram's docs, only Nova-3 with the language set to hy. Neither multi mode nor Nova-2 covers Armenian.

Why isn't Scribe's 2% a win?

The clips were voiced by ElevenLabs, and Scribe recognises them suspiciously well: 2.3% clean and 2.3% with noise louder than speech. ElevenLabs itself puts Armenian in its 5–10% error band — the vendor's own figure, for the newer Scribe v2.

How we measured
  • Audio: one Armenian text (4 sentences, 44 words, about library opening hours), written for this test. Voiced in ElevenLabs with the voice Sarah, model eleven_v3, 23 seconds long. Two more clips were made from it by a script: white noise with speech 8 dB louder than the noise; and noise 3 dB louder than speech, plus a narrow ringing tone. Three clips in total.
  • Models: Whisper large-v3 and large-v3-turbo on Groq, Deepgram Nova-3 (language hy and multi mode), Deepgram Nova-2, Fish Audio transcribe-1, ElevenLabs Scribe v1. Language hy was passed to every model except Nova-3 multi. The bot doesn't pass a language — the model detects it — and these clips were not run that way.
  • Metric:word error rate (WER) — words substituted, dropped or added, divided by the number of reference words. Text is lowercased and punctuation removed, nothing else: “9-ին” vs “իննին” counts as an error. Each clip was run 3 times; the table shows the median.
  • Date: 22 September 2026.
  • Caveats:three clips of one text in one synthetic voice is a small sample — read it as a direction, not a ranking. Real speech with accents and street noise may give different numbers. Scribe is hearing its own company's voice, so its row is not comparable. Code, audio and raw responses are in the stt-benchmarks repository.