# Speech recognition evaluation worksheet Original planning resource. Companion CSV sentences are synthetic, not recorded audio or measured results. Expand them with representative consented speech; a 12-prompt pack is not a sufficient production benchmark. ## Freeze the test specification - Workflow and consequential errors: - Service/model/version, date, settings, and dictionary version: - Development sessions and held-out sessions: - Vocabulary available before each test recording: freeze the list before scoring. Keep any experiment supplied with the correct answer terms separate from realistic production hints; do not use the corrected reference to build the held-out list. - Languages, accents, microphones, distance, noise, and code-switch coverage: - Reference reviewers and uncertainty policy: - Allowed normalization (digits, punctuation, identifiers): - Rejection criteria and reviewer workflow: ## One row per utterance | Audio ID | Split | Condition | Raw reference | Normalized reference | Raw output | Normalized output | S/D/I | Critical occurrences/correct | Identifiers/correct | False critical substitutions | False critical insertions | Meaning changed? | Reference certainty | Review minutes | |---|---|---|---|---|---|---|---|---|---|---|---|---|---|---| | | | | | | | | | | | | | | | | Use stable audio IDs, not personal names. Preserve source recognition, translation, and summary as separately scored stages. ## Aggregate with explicit denominators - WER = (substitutions + deletions + insertions) / reference words. - Critical-term recall = correct critical occurrences / reference critical occurrences. - Exact identifier accuracy = correct complete identifiers / reference identifier occurrences. - False critical substitution rate = critical occurrences replaced by a wrong term / reference critical occurrences. Distinguish replacements from omissions. - Negative utterance false-insertion rate = negative-control utterances with at least one false critical insertion / negative-control utterances. Also report insertion count. - Report failed/excluded audio and reference-uncertain items separately. - Report reviewer minutes and downstream meaning-changing errors. ## Fictional arithmetic check 20 word errors / 1,000 words = 2% WER. 16/20 critical occurrences = 80% recall. 8/10 identifiers = 80% accuracy. A different illustrative configuration gives 19/20 = 95% recall, with 2/100 negative-control utterances affected by false insertions. These are invented numbers, not test outcomes. ## Decision - Measured criteria met: - Unresolved failures: - Configuration accepted/revised/rejected: - Owner, date, and retest trigger: