Transcription and on-device AI

Word error rate, explained for people choosing a transcription tool

What word error rate counts, why vendors' accuracy figures are hard to compare, and how to test a transcription tool on your own meetings.

By the Notey team at AInject · · 6 min read

In short

Word error rate (WER) is the standard measure of how far a transcript is from what was actually said. It counts the words a recogniser got wrong and divides by the number of words spoken. Lower is better.

It is a useful number, and a slippery one: two WER figures are only comparable if they were measured on the same audio and scored the same way, which is rarely true of the figures on vendors' websites.

This guide explains what WER counts, what it misses, why published benchmarks are hard to compare, and how to test tools on your own meetings.

What word error rate counts

To measure WER you need two transcripts of the same audio: a reference, written carefully by a person, and the hypothesis, the recogniser's output. The two are lined up word by word, and each difference is one of three kinds:

  • Substitution (S): a word was replaced by another. "Fifteen" became "fifty".
  • Deletion (D): a word that was said is missing.
  • Insertion (I): a word appears that was never said.

With N the number of words in the reference:

WER = (S + D + I) / N

Take the reference "we can ship on Thursday", five words. If the transcript reads "we can't ship Thursday", "can" became "can't" (one substitution) and "on" is missing (one deletion). WER is 2 / 5, or 40%.

Because insertions do not add to N, a transcript full of words that were never said can score above 100%. It is not a percentage of something in the usual sense.

What it does not tell you

Every word weighs the same. In the example above, losing "on" and turning "can" into "can't" count equally, but only one of them reverses the decision. A transcript that gets every filler word right and one price wrong can beat a transcript that fumbles fillers and gets the price right.

It depends on normalisation. Before scoring, both transcripts are usually normalised: lower case, no punctuation, numbers written one way, fillers such as "uh" removed. Change those rules and the same transcript scores differently. "15" against "fifteen" is an error under one set of rules and a match under another.

It does not measure readability. Punctuation, paragraphs and capitalisation are usually stripped before scoring, so two transcripts with the same WER can be very different to read.

It says nothing about who spoke. Attributing lines to people is a separate problem with its own measures; see speaker diarization.

Why vendor benchmarks are hard to compare

A WER figure without its test set is close to meaningless. Several things vary from one published number to the next.

The audio

Clear, read speech from one speaker is far easier than a meeting with several people, crosstalk, compression from the call app and a range of accents. The same engine can score far better on audiobook-style recordings than on real calls. When a vendor quotes a figure, the first question is: on what?

Overlap with training data

Many public test sets are old and widely used. If a model was trained on audio that is in the test set, or very like it, it will score better there than on audio it has never met. The Earnings-22 dataset, 119 hours of English earnings calls from companies around the world, was published in 2022 partly to give a harder test drawn from real, accented business speech (Del Rio et al., 2022, checked 25 September 2026). Even so, any public test set can find its way into a later model's training data.

The scoring rules

Normalisation, whether fillers count, and how numbers are written all move the result. The Hugging Face Open ASR Leaderboard exists to make this consistent: it runs many open and commercial systems on the same datasets with one text normalisation, closely following Whisper's English normaliser, and reports WER alongside speed (Srivastav et al., "Open ASR Leaderboard", checked 25 September 2026). A leaderboard like that is the fairest public comparison there is, and it still measures its test sets, not your meetings.

Live or after the fact

A recogniser working live has to commit to words with little of what comes next. Given the whole recording, it can use context in both directions. Published figures are often for the easier, whole-file case. See live transcription vs transcribing after the meeting.

How to test a tool on your own audio

If transcription accuracy matters to your choice, an afternoon of testing will tell you more than any published figure.

  1. Pick representative audio. Three or four short extracts, a few minutes each, from real meetings of the kind you have: your usual microphone, your usual call app, the accents and jargon of your work. Get permission from the people in them to use the recordings this way.
  2. Write a reference. Type exactly what was said, listening as many times as you need. This is the slow part and the part that makes the test worth anything.
  3. Run each tool on the same audio. Same files, same settings.
  4. Normalise both sides the same way. Lower case, strip punctuation, and pick one way to write numbers.
  5. Score. A spreadsheet will do for short extracts, or use an open-source WER library such as jiwer for Python.
  6. Then read the errors. Mark the names, numbers and negations each tool got wrong. For meetings, that list matters more than the headline figure.

Keep the extracts short and the reference careful. A sloppy reference measures your typing, not the recogniser.

What this means for choosing a notetaker

WER is worth knowing about, but for meeting notes it is rarely the deciding factor on its own. The audio you feed a recogniser changes its WER more than switching between good engines does, and the errors that matter are a handful of names and figures that any engine can get wrong. Why meeting transcripts get names and numbers wrong goes through those, and getting better meeting transcripts covers the fixes on your side of the microphone.

Other questions often decide it instead: where the audio is transcribed, whether a recording is kept so you can check a line, and whether the tool works with no network. On-device transcription covers what running the recogniser on your own Mac changes.

Frequently asked questions

What is a good word error rate?

There is no single threshold, because the figure depends on the audio. The same engine scores far better on clear read speech than on a noisy meeting. Compare engines only on the same audio, scored the same way.

How is word error rate calculated?

Add the words the transcript substituted, deleted and inserted compared with a correct reference transcript, and divide by the number of words in the reference. Multiply by 100 for a percentage.

Can word error rate be more than 100%?

Yes. Insertions count as errors but do not add to the reference length, so a transcript that adds many words that were never said can score above 100%.

Is accuracy the same as 100% minus the word error rate?

Vendors often present it that way, but it is a simplification. It hides which kind of errors were made and says nothing about which words were wrong.

How do I test a transcription tool myself?

Take a few minutes of your own meetings, type a careful reference transcript, run the same audio through each tool, and score them the same way. Then look at the names and numbers each one got wrong.