How to test dictation accuracy without fooling yourself
A fair dictation test uses the same audio, reference transcript, settings, and environment for every system. Word error rate is useful, but a buying decision also needs correction time, punctuation, names, latency, failure behavior, and the effort required to put usable text where you work.
Last verified: 2026-07-31. Voicetypr publishes this protocol and benefits if readers consider its product. We cite established scoring documentation but report no Voicetypr accuracy result because we did not run a controlled study for this page. This is a test design, not a benchmark.
Quick verdict
Use a human-checked reference transcript and calculate substitutions, deletions, and insertions for a basic word error rate. Keep raw model output separate from automatic cleanup. Add a correction-effort score and workflow timing, because a low WER can still hide wrong names, punctuation, or slow insertion.
Verdict by role
Individual buyer
Use a short paired workflow test
Three representative passages and correction time are more useful than a vendor’s unrelated benchmark.
Researcher or team evaluator
Pre-register a larger test set
Balanced speakers, devices, languages, and conditions reduce cherry-picking and make results reproducible.
Accessibility user
Measure total interaction burden
Correction gestures, hotkey reach, fatigue, and failed fields may matter more than raw WER.
Decision criteria
Identical input
Use the same recorded audio when comparing model output. Live repetition changes pace, pronunciation, and noise.
Reference quality
Create a human-verified transcript with documented rules for punctuation, numbers, fillers, casing, and proper names.
Raw versus cleaned output
Score raw speech recognition separately from punctuation, rewriting, or LLM formatting.
Workflow cost
Track capture time, processing delay, corrections, cursor errors, and the time until text is usable in the destination.
| Metric | What it catches | How to record | Blind spot |
|---|---|---|---|
| WER | Inserted, deleted, substituted words | (I + D + S) / reference words | Punctuation and severity |
| Correction time | Real cleanup burden | Seconds to approved text | Reviewer skill |
| Critical-term errors | Names, numbers, negation, jargon | Count predefined critical tokens | Not a full-language score |
| End-to-end latency | Workflow delay | Speech end to usable insertion | Does not measure quality |
Build the test set before choosing a winner
Include short messages, a long paragraph, names, numbers, abbreviations, technical terms, and at least one difficult acoustic condition you genuinely face. Record clean audio once, then preserve it with the reference transcript and settings.
Do not build the test from sentences a specific model already handles well. If multiple people will use the tool, include them or state clearly that the result applies to one speaker.
Calculate WER, then inspect the errors
Microsoft describes word error rate as insertions plus deletions plus substitutions divided by reference words. NIST’s SCTK provides the `sclite` scoring tool for reproducible comparison.
A single percentage is not enough. A substituted project name or deleted “not” can cost more than several harmless filler errors. Label critical terms and publish examples alongside the aggregate.
Measure the text users actually receive
Save raw recognition output before any AI rewrite. If a product applies cleanup, score that output separately and record whether it changes meaning, numbers, uncertainty, or quoted language.
Time the complete path from the end of speech to usable text in the target field. Include manual copying, failed pastes, formatting repair, and correction actions.
Report limits so others can interpret the result
Publish model versions, hardware, microphone, language, speakers, room conditions, decoding settings, and test date. A score without this context should not be treated as a product-wide fact.
Repeat uncertain cases rather than deleting them. If the test is too small for statistical claims, call it a personal workflow trial. Honest scope is more useful than false precision.
Limitations and checks
- WER depends on normalization and reference-transcript rules.
- A small personal test cannot establish population-wide accuracy.
- Live-dictation usability includes factors WER does not measure.
- This page supplies no product ranking or measured result.
How we evaluated
- Based the core error metric on Microsoft’s current evaluation guide and NIST SCTK.
- Separated recognition output, post-processing, and user correction.
- Added workflow and critical-term measures for buyer decisions.
- Made all absent testing explicit rather than implying a laboratory result.
Sources
- Microsoft: test speech-model accuracy
- NIST: Speech Recognition Scoring Toolkit
- OpenAI: Whisper research paper
- Microsoft: Speech-to-text FAQ
- Voicetypr public desktop repository
Recheck pricing, requirements, and privacy terms with each provider before buying.
Frequently asked questions
What is a good word error rate?
There is no universal buying threshold. Compare systems on the same representative set, inspect important errors, and include correction effort and latency.
Can I compare by reading the same paragraph twice?
You can run a personal trial, but it is not controlled because pace, pronunciation, and noise change. Recorded audio is better for model comparison.
Should punctuation count?
Document it separately unless your normalization rules include it. Punctuation can materially affect usability even when lexical WER is unchanged.
Run a transparent three-passage trial
Use the free trial with saved audio, keep raw output, and record correction time in the applications that matter to you.