“ZX-204-B” becomes “ZX-204-D,” and a fluent meeting summary points purchasing to the wrong component. Adding a dictionary entry helps only if recognition is where the error began.
Trace ten recent identifier errors from recording to final request. Label the first stage that changed each one. This small diagnostic set helps choose the repair; it is not an accuracy benchmark.
Find the first changed character
When the source is audible but an ordinary word replaces the term, investigate recognition vocabulary or model fit. If the words are right and only the digits or separators change, inspect formatting. If a correct source becomes a wrong translated unit, investigate translation. If “could review” becomes “accepted,” investigate the summary and approval step.
Keep raw and normalized text. An agreed rule may equate “zero point eight millimeters” with “0.8 mm”; it must preserve the decimal and unit. Mark an inaudible reference uncertain instead of scoring a plausible guess as correct. Microsoft evaluation guidance .
Match the vocabulary control to its owner
| Option | Vocabulary control and limit |
|---|---|
| Azure Custom Speech | Maintain domain data, evaluation and a custom endpoint. Supported data types, pronunciation and display controls vary by language/configuration. |
| Google adaptation | Maintain phrase sets and contextual bias. Stronger bias can insert a phrase never spoken. |
| Deepgram Keyterm Prompting | Nova-3/Flux: plain terms, 500 tokens total per request. No legacy Keywords weights/intensifiers. |
| VoicePing dictionary | Paid personal/workspace transcription and translation entries. These controls serve different stages. |
Sources: Azure , Google , Deepgram , VoicePing .
A team building a speech application can compare API controls and model ownership. A team improving recurring meetings should first check its existing application’s terminology controls. For VoicePing, export existing dictionary settings before an overwrite.
Score the fields people act on
In a fictional 1,000-word test, 20 word errors give 2% word error rate, while only eight of ten complete identifiers are correct: 80% exact identifier accuracy. The first number cannot substitute for the second.
Use the synthetic terminology pack and scoring worksheet with representative recordings and a domain reviewer’s reference. Include similar identifiers, quantities, negation and language switches. Freeze the vocabulary before testing fresh speakers or sessions.
After tuning, also test ordinary utterances that contain none of the target terms. Record false insertions against that separate negative-control set, alongside exact identifier matches and review time. A dictionary that recovers a part number but invents it elsewhere has introduced another problem to resolve.
Documentation reviewed September 6, 2026. Examples are synthetic; no comparative product results are claimed.






