How to test dictation accuracy without fooling yourself

A fair dictation test uses the same audio, reference transcript, settings, and environment for every system. Word error rate is useful, but a buying decision also needs correction time, punctuation, names, latency, failure behavior, and the effort required to put usable text where you work.

Quick verdict

Use a human-checked reference transcript and calculate substitutions, deletions, and insertions for a basic word error rate. Keep raw model output separate from automatic cleanup. Add a correction-effort score and workflow timing, because a low WER can still hide wrong names, punctuation, or slow insertion.

Verdict by role

Individual buyer

Use a short paired workflow test

Three representative passages and correction time are more useful than a vendor’s unrelated benchmark.

Researcher or team evaluator

Pre-register a larger test set

Balanced speakers, devices, languages, and conditions reduce cherry-picking and make results reproducible.

Accessibility user

Measure total interaction burden

Correction gestures, hotkey reach, fatigue, and failed fields may matter more than raw WER.

Decision criteria

Identical input

Use the same recorded audio when comparing model output. Live repetition changes pace, pronunciation, and noise.

Reference quality

Create a human-verified transcript with documented rules for punctuation, numbers, fillers, casing, and proper names.

Raw versus cleaned output

Score raw speech recognition separately from punctuation, rewriting, or LLM formatting.

Workflow cost

Track capture time, processing delay, corrections, cursor errors, and the time until text is usable in the destination.

Metrics for a useful dictation test
MetricWhat it catchesHow to recordBlind spot
WERInserted, deleted, substituted words(I + D + S) / reference wordsPunctuation and severity
Correction timeReal cleanup burdenSeconds to approved textReviewer skill
Critical-term errorsNames, numbers, negation, jargonCount predefined critical tokensNot a full-language score
End-to-end latencyWorkflow delaySpeech end to usable insertionDoes not measure quality

Build the test set before choosing a winner

Include short messages, a long paragraph, names, numbers, abbreviations, technical terms, and at least one difficult acoustic condition you genuinely face. Record clean audio once, then preserve it with the reference transcript and settings.

Do not build the test from sentences a specific model already handles well. If multiple people will use the tool, include them or state clearly that the result applies to one speaker.

Calculate WER, then inspect the errors

Microsoft describes word error rate as insertions plus deletions plus substitutions divided by reference words. NIST’s SCTK provides the `sclite` scoring tool for reproducible comparison.

A single percentage is not enough. A substituted project name or deleted “not” can cost more than several harmless filler errors. Label critical terms and publish examples alongside the aggregate.

Measure the text users actually receive

Save raw recognition output before any AI rewrite. If a product applies cleanup, score that output separately and record whether it changes meaning, numbers, uncertainty, or quoted language.

Time the complete path from the end of speech to usable text in the target field. Include manual copying, failed pastes, formatting repair, and correction actions.

Report limits so others can interpret the result

Publish model versions, hardware, microphone, language, speakers, room conditions, decoding settings, and test date. A score without this context should not be treated as a product-wide fact.

Repeat uncertain cases rather than deleting them. If the test is too small for statistical claims, call it a personal workflow trial. Honest scope is more useful than false precision.

Limitations and checks

  • WER depends on normalization and reference-transcript rules.
  • A small personal test cannot establish population-wide accuracy.
  • Live-dictation usability includes factors WER does not measure.
  • This page supplies no product ranking or measured result.

How we evaluated

  1. Based the core error metric on Microsoft’s current evaluation guide and NIST SCTK.
  2. Separated recognition output, post-processing, and user correction.
  3. Added workflow and critical-term measures for buyer decisions.
  4. Made all absent testing explicit rather than implying a laboratory result.

Frequently asked questions

What is a good word error rate?

There is no universal buying threshold. Compare systems on the same representative set, inspect important errors, and include correction effort and latency.

Can I compare by reading the same paragraph twice?

You can run a personal trial, but it is not controlled because pace, pronunciation, and noise change. Recorded audio is better for model comparison.

Should punctuation count?

Document it separately unless your normalization rules include it. Punctuation can materially affect usability even when lexical WER is unchanged.

Congrats! 🎉

Your purchase was successful.

You will receive an email with your purchase details.