Speech to Text Accuracy: What WER Means and How to Test It Yourself
Speech to text accuracy is almost always measured with word error rate (WER): the number of words a system gets wrong, divided by the number of words actually spoken. Lower is better, so when an app says it's "95% accurate", it usually means roughly a 5% WER on audio the vendor picked, which may sound nothing like you at your desk.
That last part is why accuracy claims are so hard to compare. Below is how WER works, what pushes it up or down, and how to measure AI transcription accuracy on your own voice in about half an hour.
What word error rate actually measures
WER compares two pieces of text. One is a reference transcript: what was really said, written down carefully by a person. The other is what the speech-to-text system produced. You line them up word by word and count three kinds of mistake:
- Substitutions (S): a word replaced with the wrong one, like "fixed" for "fix".
- Deletions (D): a spoken word that never made it into the text.
- Insertions (I): words the system added that nobody said.
Add those up and divide by the number of words in the reference (N):
WER = (S + D + I) / N
Microsoft's Speech service docs call WER the industry standard for measuring model accuracy, and use exactly this formula.
A worked example
Say you dictate this nine-word sentence:
push the fix to the staging branch before lunch
And the app types:
push the fixed to staging branch before the lunch
That's one substitution ("fixed"), one deletion (the first "the" before "staging") and one insertion (the extra "the" before "lunch"). Three errors over nine words gives a WER of 33%.
Read the output again, though. You understood it fine. That's the odd thing about WER: a transcript can be perfectly readable and still score badly, and a transcript with one catastrophic mistake (a wrong name, a wrong number) can score well.
Why jargon gets punished twice
Here's a second sentence with the same score:
rename the user id field in the postgres schema
If the app writes "rename the user idea field in the post gress schema", you get a substitution for "idea", then "postgres" becomes two words: another substitution plus an insertion. Three errors, 33% again, and two of them come from one technical term. If you work with a lot of product names or code terms, your personal WER will look worse than any benchmark.
WER can go over 100%
Because you divide by the length of the reference, an app that hallucinates a lot of extra words can rack up more errors than there were words spoken. Wikipedia's entry on word error rate notes that WER can be larger than 1.0 when insertions outnumber correct words. Word accuracy is sometimes quoted instead, calculated as 1 minus WER, which is why it can technically go below zero.
What counts as good speech to text accuracy?
Microsoft's guidance is a useful yardstick. In the same docs, it says a WER of 5 to 10% is good quality and ready to use, 20% is acceptable but might need more training, and 30% or more signals poor quality. Its scenario table puts dictation among the cases that can get under 10% WER, while call-centre audio sits under 30%.
Now do the maths on something you'd actually write. A 200-word email at 5% WER has about ten word errors in it. That counts as good quality by Microsoft's yardstick. It's also ten things to spot and fix before you press send. So for dictation, the number that matters to you is less "how accurate is it" and more "how long do I spend correcting it".
Why vendor accuracy numbers don't match your experience
Every accuracy figure comes from a test set, and test sets vary wildly. Some are clean read speech. Others are messy, like earnings calls with crosstalk and financial jargon. Hugging Face's Open ASR Leaderboard scores models on a whole spread of public datasets, from short English clips to long-form earnings calls and multilingual sets, and a model that tops one column can look ordinary in the next.
Normalisation matters too. Standard WER is calculated on the plain words, so "Hello, John." and "hello john" count as a match. For dictation, though, punctuation and casing are part of what you have to fix. Microsoft offers a separate measure, token error rate, that keeps punctuation, capitalisation and formatting (like writing "$900" instead of "nine hundred dollars") in the score.
So when you see "99% accurate", ask on what audio, scored how. If the page doesn't say, treat the number as marketing.
What affects speech recognition accuracy
Some of this you can control. Some of it you can't.
Your microphone and how far away it is
Microsoft's docs link lots of deletion errors to weak audio signal and suggest recording closer to the source. A laptop mic across the desk picks up the room as much as your voice. A headset or a decent USB mic a hand's width from your mouth is usually the single biggest fix.
Background noise and other voices
The same docs tie insertion errors to noisy environments and crosstalk. A TV in the next room, a colleague on a call, a coffee machine: the model tries to transcribe all of it.
Jargon, names and acronyms
Substitution errors, per Microsoft, often happen when a model hasn't seen enough of your domain's terms. This is where custom vocabulary or a personal dictionary pays off.
Accent and speaking style
How much an accent or dialect affects accuracy depends heavily on what audio the model was trained on. Speaking style matters too. Careful dictation and a rambling phone call are very different inputs, and Microsoft's scenario table treats them as separate cases with separate expectations.
Audio quality
Microsoft's scenario table lists call-centre audio at 8 kHz and dictation at 16 kHz, and the gap shows up in the error rates. Bluetooth headsets are a quiet culprit here, since using their mic often drops them into a lower-quality call mode without you noticing.
The model itself
Different models really do score differently, and the same model can rank well on one dataset and badly on another. Where the model runs (on your machine or in the cloud) tells you about privacy and latency. It doesn't tell you the accuracy; only a test does.
If your accuracy suddenly gets worse rather than being consistently mediocre, it's usually a setup problem rather than the model. Our speech to text troubleshooting guide covers the common causes, such as the wrong input device being selected.
How to test speech to text accuracy yourself
You don't need a lab. You need a script, a quiet-ish room and a bit of Python.
- Write a reference script of 200 to 300 words. Make it look like what you actually dictate: a work email, a Slack reply, a prompt to an AI agent. Include the names, product terms and acronyms you use every day.
- Record or dictate it in your normal setup. Same mic, same room, same speaking pace. Don't do your newsreader voice unless that's how you'll really talk to it.
- Run it through each app you're comparing. For live dictation apps, read the script into each one and save the output to a text file. For file-based tools, record once and feed every tool the same audio, which is the fairer test.
- Score it. The open-source jiwer library for Python does the alignment for you. Install it with
pip install jiwer, save your script asreference.txtand an app's output asapp-a.txt, then run:
import jiwer
normalise = jiwer.Compose([
jiwer.ToLowerCase(),
jiwer.RemovePunctuation(),
jiwer.RemoveMultipleSpaces(),
jiwer.Strip(),
jiwer.ReduceToListOfListOfWords(),
])
reference = open("reference.txt").read()
hypothesis = open("app-a.txt").read()
out = jiwer.process_words(
reference,
hypothesis,
reference_transform=normalise,
hypothesis_transform=normalise,
)
print("WER:", round(out.wer * 100, 1), "%")
print("Substitutions:", out.substitutions)
print("Deletions:", out.deletions)
print("Insertions:", out.insertions)
- Look at the error types, not just the total. A pile of deletions points to your mic or distance. Insertions usually mean noise. And if the substitutions keep landing on the same few words, you need a custom dictionary, or a different app.
- If you've got the patience, repeat it with the kettle on or on AirPods instead of a desk mic, because real life is rarely the quiet run.
If Python isn't your thing, do it by hand: print both texts, circle every wrong, missing and extra word, and divide by the word count of your script. Tedious, but it works.
A more practical measure for daily dictation is corrections per 100 words. Dictate normally for a week and keep a tally of how often you go back to fix something. It captures punctuation, formatting and the misses that actually annoy you, which a raw WER doesn't.
Where Blazing Transcribe fits
We make Blazing Transcribe, a free dictation app for Mac, so take this with that in mind. Its homepage makes a speed claim (155× faster than real time) and doesn't put an accuracy percentage on it, and after everything above you can probably see why: your number depends on your mic, your room, your accent and your vocabulary.
What we'd suggest instead is the test in this article. Blazing runs on-device, has a personal dictionary for the names and jargon that cause substitutions, and costs nothing, so trying it against whatever you use now costs you half an hour. For the wider field, our roundups of the best speech to text software and AI dictation software are a good place to find the other apps worth testing.
FAQ
What is a good accuracy rate for speech to text?
By Microsoft's guidance, a word error rate of 5 to 10% is good quality, which is roughly 90 to 95% word accuracy. Around 20% is usable but rough, and 30% or more is poor. For dictation you edit yourself, aim for the low end and check how much time you spend on corrections.
How is AI transcription accuracy measured?
With word error rate. You compare the AI's transcript against a careful human transcript, count substitutions, deletions and insertions, and divide by the number of words in the human version. Many tools strip punctuation and capitals before scoring, so check how the number was calculated.
Is 95% accurate transcription good?
It's decent, but it means about one word in twenty is wrong. On a 1,000-word document that's around 50 errors to find. Whether that's good enough depends on whether you're skimming notes or sending something to a client.
Why is my speech to text so inaccurate?
The most common causes are a distant or low-quality mic, background noise, and vocabulary the model doesn't know. Check which input device is selected, get closer to the mic, and add your common names and terms to a custom dictionary if your app supports one.
Is WER the same as accuracy?
They're two sides of one number. Word accuracy is usually reported as 1 minus WER, so a 7% WER is about 93% accuracy. Because WER can go above 100% when an app inserts lots of extra words, accuracy can technically go below zero.