
In this article
Today we are introducing VoicePing ASR Model V0.1, our speech-to-text model focused on the Asian languages that matter most to VoicePing: Japanese, Korean, Chinese, and Vietnamese, alongside English.
VoicePing is built around spoken communication across Asia: meetings, events, voice translation, transcripts, summaries, and search. In those workflows, ASR is not an isolated feature. It is the first layer of the entire product experience. If the transcript is unstable, every downstream step becomes less useful.
VoicePing ASR Model V0.1 is designed for that reality. It focuses on Japanese, Korean, Chinese, Vietnamese, and English, with the goal of producing cleaner transcripts for real conversations.
One Model For Asian Languages And English
General-purpose speech recognition has improved quickly, but real speech across Asian languages and English still has difficult edges:
- Japanese, Korean, and Chinese need language-aware text handling.
- Vietnamese depends on accurate tone marks and word boundaries.
- Long or noisy clips can expose partial transcripts, empty outputs, and repeated text.
- Cloud models can behave differently across languages, even when the API looks uniform.
- A system that performs well on a public benchmark is not always the best fit for meetings, events, and voice translation.
VoicePing ASR Model V0.1 is our first consolidated model built around this Asian-language product surface. The benchmark below asks a practical question: how well does it transcribe the speech our users actually care about?
What It Does
VoicePing ASR Model V0.1 transcribes speech in:
- English
- Japanese
- Korean
- Chinese
- Vietnamese
The output is the transcript that powers later VoicePing features such as translation, captions, meeting notes, and searchable conversation history.
This article uses ASR and STT interchangeably. Both mean speech-to-text transcription.
Evaluations
Dataset
The evaluation uses an Asian-language-focused VoicePing speech set with 1,000 clips per language, about 41 hours of audio in total. The clips reflect the kind of speech VoicePing handles in practice: real conversations rather than clean read-aloud recordings.
| Language | Clips |
|---|---|
| English | 1,000 |
| Japanese | 1,000 |
| Vietnamese | 1,000 |
| Korean | 1,000 |
| Chinese | 1,000 |
| Total | 5,000 |
Every system is tested on the same audio set.
Models Compared
We compare VoicePing ASR Model V0.1 with widely used cloud speech systems and open ASR models, including Google Cloud STT, Azure AI Speech, OpenAI transcription models, ElevenLabs Scribe v2, Deepgram Nova-3, Qwen3-ASR, and SenseVoiceSmall.
Scoring
The headline metric is word error rate (WER): lower is better. WER measures how many words are inserted, deleted, or substituted compared with the human reference transcript.
WER is an edit count divided by the number of reference words, not a sentence-pass percentage. Tokenization and normalization matter, especially when word boundaries, scripts, punctuation or numerals differ. The public article does not provide the complete language-specific scoring code, so its values should be compared within this reported run rather than treated as interchangeable with another benchmark’s WER. Macro WER then averages the five language results with equal weight; it does not weight languages by a customer’s traffic or by total audio duration.
Latency
Accuracy is not the only requirement for production ASR. We also measure how long each system takes to return a transcript, because a model that is accurate but slow can still feel poor in live meetings and events.
The table reports observed median response time. It does not separate time to first partial text, time to stable text and time to the final transcript, nor does it normalize local and API serving paths. A live-captioning decision needs those timing boundaries and behavior under load. A short median by itself does not establish streaming responsiveness or complete output.
Main Results
The chart below compares the average word error rate across the five languages for VoicePing ASR Model V0.1 and the external speech-to-text systems in this benchmark. Lower bars are better. Per-language charts follow in the results section.

Accuracy by Language
| System | EN WER | JA WER | VI WER | KO WER | ZH WER | Macro WER |
|---|---|---|---|---|---|---|
| VoicePing ASR Model V0.1 | 20.2% | 20.4% | 15.5% | 24.5% | 16.0% | 19.3% |
| Google Cloud STT V1 default | 23.1% | 23.5% | 52.1% | 57.8% | 44.2% | 40.1% |
| Google Cloud STT Chirp 2 | 24.5% | 29.7% | 14.8% | 32.8% | 22.6% | 24.9% |
| Google Cloud STT Chirp 3 | 22.9% | 26.4% | 20.1% | 37.4% | 19.2% | 25.2% |
| Azure AI Speech | 23.0% | 21.1% | 21.0% | 37.3% | 22.5% | 25.0% |
| OpenAI GPT-4o Transcribe | 50.6% | 52.4% | 64.3% | 44.1% | 29.1% | 48.1% |
| OpenAI GPT Realtime Whisper | 31.8% | 26.2% | 20.9% | 33.0% | 22.4% | 26.9% |
| Qwen3-ASR 0.6B | 23.8% | 29.7% | 26.2% | 38.2% | 20.9% | 27.7% |
| Qwen3-ASR 1.7B | 21.9% | 25.0% | 22.0% | 33.1% | 20.0% | 24.4% |
| SenseVoiceSmall | 28.0% | 37.4% | 99.9% | 45.9% | 28.1% | 47.9% |
| ElevenLabs Scribe v2 | 28.6% | 20.3% | 15.4% | 31.5% | 21.2% | 23.4% |
| Deepgram Nova-3 | 29.3% | 28.0% | 38.4% | 44.8% | 29.2% | 34.0% |
Leaderboard
| System | Macro WER | Median latency | Notes |
|---|---|---|---|
| VoicePing ASR Model V0.1 | 19.3% | 1.22s | VoicePing Asian-language ASR |
| Google Cloud STT V1 default | 40.1% | 7.47s | Cloud speech-to-text |
| Google Cloud STT Chirp 2 | 24.9% | 7.12s | Cloud speech-to-text |
| Google Cloud STT Chirp 3 | 25.2% | 7.32s | Cloud speech-to-text |
| Azure AI Speech | 25.0% | 7.12s | Cloud speech-to-text |
| OpenAI GPT-4o Transcribe | 48.1% | 1.53s | OpenAI transcription |
| OpenAI GPT Realtime Whisper | 26.9% | 7.17s | OpenAI transcription |
| Qwen3-ASR 0.6B | 27.7% | 3.56s | Open ASR model |
| Qwen3-ASR 1.7B | 24.4% | 4.18s | Open ASR model |
| SenseVoiceSmall | 47.9% | 0.07s | Open ASR model |
| ElevenLabs Scribe v2 | 23.4% | 3.07s | Cloud speech-to-text |
| Deepgram Nova-3 | 34.0% | 1.33s | Cloud speech-to-text |
VoicePing ASR Model V0.1 combines the lowest macro WER in this benchmark with one of the fastest median response times, at 1.22 seconds. SenseVoiceSmall returned a shorter median, but its five-language score includes Vietnamese outside its documented language set. This is not an equivalent coverage comparison, and the table does not establish that faster execution caused lower accuracy.
System-by-system results
The twelve entries follow the published table order, not a performance ranking. Macro WER gives each of the five languages equal weight; lower is better. Every row uses the same 5,000-clip test set, while request paths and model language coverage differ. These summaries restate the original benchmark, not a new evaluation.
1. VoicePing ASR Model V0.1
Macro WER: 19.3% · Observed median latency: 1.22s.
VoicePing recorded the lowest macro WER in this five-language set. Its English, Korean and Chinese rows were 20.2%, 24.5% and 16.0%. Japanese was 20.4%, close to Scribe v2’s 20.3%, while Chirp 2 led Vietnamese at 14.8% against VoicePing’s 15.5%. The aggregate therefore supports a strong result on this corpus without establishing a lead in every language.
For a product decision, the next question is which remaining errors affect names, numbers and meaning. A 19.3% WER is an edit-distance ratio, not a guarantee that a particular fraction of sentences is usable. The 1.22-second median describes the tested transcription response, not the complete microphone-to-caption or speech-translation path. This publication does not measure downstream summary accuracy, streaming revision stability or concurrent capacity. Those need separate acceptance criteria before extending this model-level result to the whole VoicePing experience.
2. Google Cloud STT V1 default
Macro WER: 40.1% · Observed median latency: 7.47s.
This is the legacy V1 default row, not a combined score for Google Cloud Speech-to-Text. English and Japanese WER were 23.1% and 23.5%, but Vietnamese reached 52.1%, Korean 57.8% and Chinese 44.2%. Because the macro average weights each language equally, those three rows have a substantial effect on the 40.1% result.
Keep the V1 request configuration separate from the Chirp 2 and Chirp 3 experiments. Google’s model-selection documentation distinguishes models and recognition methods; a provider name alone does not define a reproducible test. The public article does not supply every locale, audio-decoding and endpoint setting, so the score cannot diagnose which factor caused the gap. A useful migration check would rerun representative audio with an explicitly selected model and method. It should not assume either that the legacy score describes current Google offerings or that migration automatically produces a measured improvement.
3. Google Cloud STT Chirp 2
Macro WER: 24.9% · Observed median latency: 7.12s.
Chirp 2 is a V2 model with recognition methods whose language coverage differs. Its Vietnamese WER of 14.8% was the lowest point estimate in this table. Korean WER was 32.8% and Japanese 29.7%, so the Vietnamese lead did not translate into the lowest five-language average.
This makes Chirp 2 relevant to a Vietnamese-heavy evaluation even when the overall ranking favors another system. The observed lead over Scribe v2 was 0.6 percentage points; the article does not publish paired uncertainty estimates establishing that small difference. Its 7.12-second median also needs the request path attached to it. A batch completion time does not tell a reader when the first stable streaming words would appear. Verify method, region and target-language support together before using this row to select a live-captioning configuration.
4. Google Cloud STT Chirp 3
Macro WER: 25.2% · Observed median latency: 7.32s.
Chirp 3 had lower WER than the tested Chirp 2 row for English, Japanese and Chinese: 22.9%, 26.4% and 19.2%. Chirp 2 was lower for Vietnamese and Korean. This mixed result explains why a newer generation need not lead every row of one product-specific dataset.
The macro difference between Chirp 3 and Chirp 2 was only 0.3 percentage points, and the publication provides no significance test for it. Read the five language rows before deciding which configuration to investigate. Google’s documentation also covers features such as diarization and language detection, but this transcript benchmark does not score their quality. The 7.32-second median is a historical observation for the tested request path. It should not be reused as a universal streaming-delay or current-service performance claim.
5. Azure AI Speech
Macro WER: 25.0% · Observed median latency: 7.12s.
Azure AI Speech had Japanese WER of 21.1%, comparatively close to the lowest two Japanese point estimates, while Korean WER was 37.3%. English, Vietnamese and Chinese were 23.0%, 21.0% and 22.5%. Those differences make a language-specific shortlist more useful than treating 25.0% macro WER as a uniform experience.
Microsoft distinguishes real-time, fast and batch transcription workflows and offers customization options. The generic service label in this export does not establish that all those paths or a custom model were tested. Its 7.12-second median cannot identify whether model compute, transfer or another serving step dominated. For an existing Azure workflow, repeat the intended language and recognition method, then review domain terms and transcript completion. No improvement from customization, phrase hints or a different endpoint was measured in this article.
6. OpenAI GPT-4o Transcribe
Macro WER: 48.1% · Observed median latency: 1.53s.
GPT-4o Transcribe is an audio-to-text model. This export reported WER from 29.1% for Chinese to 64.3% for Vietnamese, producing a 48.1% macro value. The 1.53-second median is relatively short, but it does not establish that the returned text was complete or suitable for the downstream task.
The unexpectedly weak result deserves a request-and-output audit before becoming a general purchasing conclusion. Inspect the actual audio submitted, full responses, language handling and scoring normalization, and separate omissions, empty outputs and substitutions where the raw record permits. Those are diagnostic steps, not established causes of this result. The public article lacks the complete request and output package needed to resolve them. Keep the reported measurement visible while limiting it to this run; do not infer a general failure rate for the model or assume an untested configuration would fix it.
7. OpenAI GPT Realtime Whisper
Macro WER: 26.9% · Observed median latency: 7.17s.
“OpenAI GPT Realtime Whisper” is the original benchmark label, retained for traceability. The public record does not resolve it to an exact API model ID, snapshot or session configuration. OpenAI’s current realtime-transcription documentation names explicit models and distinguishes incremental transcript events from final turn completion. That documentation cannot retroactively identify which historical system produced this row.
The row’s language values ranged from 20.9% WER for Vietnamese to 33.0% for Korean. Its 7.17-second median should not be interpreted as time to first partial text or stable live captions without a documented timing boundary. To make this comparison reproducible, the original request record needs to establish model selection, turn handling and the output used for scoring. Until then, this is a partially identified historical result, not a recommendation to request a model literally named “GPT Realtime Whisper.”
8. Qwen3-ASR 0.6B
Macro WER: 27.7% · Observed median latency: 3.56s.
Qwen3-ASR 0.6B is the smaller of the two Qwen checkpoints compared. Its published language set includes all five evaluation languages. In this run, WER ranged from 20.9% for Chinese to 38.2% for Korean. The 1.7B row had lower WER in every language, so the macro difference was not driven by just one category.
The smaller checkpoint returned results sooner in this setup: 3.56 seconds at the median versus 4.18 seconds for 1.7B. This is an observed trade-off, not a controlled scaling law. The model card describes multiple inference backends and streaming support; the article does not show that every backend or mode was evaluated. For a local deployment, measure memory, concurrency and transcript completion using the intended runtime. Parameter count and a single median cannot establish the number of simultaneous users the system will support.
9. Qwen3-ASR 1.7B
Macro WER: 24.4% · Observed median latency: 4.18s.
Qwen3-ASR 1.7B improved the macro result from the smaller checkpoint’s 27.7% to 24.4%. English WER was 21.9%, Japanese 25.0%, Vietnamese 22.0%, Korean 33.1% and Chinese 20.0%. The direction was consistent across all five rows, although the article does not provide paired uncertainty estimates for each gap.
The quality improvement came with a longer observed median, 4.18 seconds against 3.56 seconds for 0.6B. Whether that trade-off is useful depends on the application: uploaded recording processing and live captions have different timing requirements. The checkpoint’s documented streaming capability does not mean this benchmark measured partial-transcript stability or endpointing delay. Keep this larger model on a local-runtime shortlist when its quality gains matter, then separately test memory, batching and load. Those production measurements cannot be recovered from the parameter label or the current latency table.
10. SenseVoiceSmall
Macro WER: 47.9% · Observed median latency: 0.07s.
The SenseVoiceSmall model card lists Mandarin, Cantonese, English, Japanese and Korean. Vietnamese is outside that stated set, yet it contributes equally to this benchmark’s macro average with a 99.9% WER result. The 47.9% macro value therefore combines a coverage mismatch with the checkpoint’s results on its documented languages.
The 0.07-second median was the shortest in the table, but it is not evidence of a usable five-language system. Even within documented coverage, English, Japanese, Korean and Chinese WER were 28.0%, 37.4%, 45.9% and 28.1%. A narrower-language application could evaluate those rows directly; a five-language application would need a different model or an explicitly tested routing design. No fallback or routed combination was measured here. Preserve the full result while making the unsupported-language condition visible instead of presenting the fast response as an equivalent replacement.
11. ElevenLabs Scribe v2
Macro WER: 23.4% · Observed median latency: 3.07s.
ElevenLabs Scribe v2 had the lowest macro WER among the external rows in this table. Japanese WER was 20.3% and Vietnamese 15.4%, close to VoicePing’s 20.4% and 15.5%. English, Korean and Chinese were higher at 28.6%, 31.5% and 21.2%, which explains why those near-ties did not produce the lowest overall average.
The small Japanese and Vietnamese differences should remain point estimates: no uncertainty analysis is supplied. The official documentation distinguishes Scribe v2 from Scribe v2 Realtime, so the recorded 3.07-second median must not be assigned to the separate live model. Scribe also offers timestamps and speaker labeling, but their accuracy is outside this WER comparison. For a meeting-transcript application, assess those outputs alongside text accuracy in a separate test, using the exact model and options intended for deployment.
12. Deepgram Nova-3
Macro WER: 34.0% · Observed median latency: 1.33s.
Deepgram Nova-3 returned a 1.33-second observed median, close to VoicePing’s 1.22 seconds in this setup. Its language results varied: Japanese WER was 28.0%, English 29.3%, Chinese 29.2%, Vietnamese 38.4% and Korean 44.8%. A single 34.0% macro average obscures which language mix drives the application decision.
The service documentation distinguishes model and language configurations, including multilingual options. This article does not publish a complete request configuration, so the row should not be generalized to every Nova-3 mode or current revision. When evaluating mixed-language meetings, save the explicit language settings and inspect switching boundaries rather than assuming the five separate language scores establish code-switching quality. Its short response time also does not score diarization, word timestamps or transcript completeness on its own. Those are additional product requirements, not outcomes measured by the latency number.
Results by Language
English
VoicePing ASR Model V0.1 has the lowest English WER in this comparison at 20.2%, ahead of Qwen3-ASR 1.7B and the major cloud systems tested here.

Qwen3-ASR 1.7B followed at 21.9%, compared with 20.2% for VoicePing. That is a corpus-level point estimate, not a claim about every English accent or recording condition. The published table has no accent-stratified results or paired confidence interval for this difference. Review errors on the microphones and speaking styles relevant to the intended application.
Japanese
Japanese is one of the most important languages for VoicePing. On this dataset, VoicePing ASR Model V0.1 reaches 20.4% WER, in a virtual tie with ElevenLabs Scribe v2 (20.3%) for the best Japanese result, ahead of Azure AI Speech and the open ASR baselines.

The 0.1-point gap between Scribe v2 and VoicePing is too small to turn into a firm ranking without an uncertainty analysis. Japanese scoring also depends on how the reference and output are segmented and normalized. For meeting use, inspect proper names, omitted particles and number rendering as well as the aggregate; this article does not publish separate counts for those error types.
Vietnamese
Vietnamese is the closest race in this benchmark: Google Cloud STT Chirp 2 leads at 14.8% WER, with ElevenLabs Scribe v2 (15.4%) and VoicePing ASR Model V0.1 (15.5%) essentially tied just behind.

The leading three point estimates fall within 0.7 percentage points. That makes error review and the intended request method important before choosing a winner. SenseVoiceSmall’s 99.9% result has a different interpretation: Vietnamese is outside the Small checkpoint’s stated language set. Keep that coverage mismatch visible rather than using the macro ranking as evidence of equivalent multilingual support.
Korean
Korean shows one of the clearest gaps in the benchmark. VoicePing ASR Model V0.1 records 24.5% WER, well ahead of the next group of systems.

The next-lowest Korean point estimate was Scribe v2 at 31.5%, a 7.0-point gap from VoicePing. This is a larger observed separation than the Japanese near-tie, but the publication does not identify its cause or provide an uncertainty interval. Spacing, endings and domain vocabulary are useful inspection categories for a follow-up; they are not measured explanations of the current gap.
Chinese
Chinese is another strong area for VoicePing ASR Model V0.1, at 16.0% WER, with Google Cloud STT Chirp 3 and Qwen3-ASR 1.7B closest behind.

Chirp 3 recorded 19.2% and Qwen3-ASR 1.7B 20.0%, compared with VoicePing’s 16.0%. The aggregate Chinese label does not establish performance for every regional accent or for Cantonese, and the article does not publish that breakdown. A deployment spanning regions should validate the intended spoken variety and text-normalization convention instead of transferring this one score to all Chinese-language audio.
What We Learned
- VoicePing ASR Model V0.1 has the strongest overall accuracy in this benchmark, with 19.3% macro WER across the five languages.
- The result is not uniform by language: English, Korean, and Chinese show the clearest VoicePing advantages in this run, Japanese is a virtual tie with the best cloud system, and Vietnamese is a close race decided by less than one point.
- Larger general-purpose systems do not automatically win on product-specific Asian-language speech.
- Accuracy and response time both matter for the user experience, especially in live meetings and events.
What’s Next
VoicePing ASR Model V0.1 is a first release, and this benchmark is a snapshot. The dataset is built from the kind of speech VoicePing handles in practice, so it provides a product-focused evaluation signal rather than establishing complete production readiness or a universal ranking — and the cloud systems in the comparison will keep evolving, as will our model. Latency also depends on the deployment environment, so treat the speed numbers as indicative rather than absolute.
From here, our work focuses on the places this evaluation points to: reducing the error patterns that remain in each language, extending the test set with noisier, longer, and more domain-specific audio, and broadening the comparison as new speech-to-text systems ship. Automatic scores guide that work, and human transcript review stays part of every release decision.
Conclusion
VoicePing ASR Model V0.1 is our first consolidated ASR model focused on Japanese, Korean, Chinese, and Vietnamese, while also supporting English for real product workflows. In this 5,000-clip benchmark it delivers the strongest overall accuracy of the systems tested, at some of the fastest response times — and it is the transcription layer every other VoicePing feature builds on.
The important shift is focus: we are evaluating ASR as part of a real communication product for Asian languages and English, not as an isolated model demo.
References
- Google Cloud STT model selection
- Azure Speech to Text
- OpenAI GPT-4o Transcribe
- OpenAI realtime transcription
- OpenAI Whisper
- ElevenLabs Scribe
- Deepgram models and languages
- Qwen3-ASR 0.6B
- Qwen3-ASR 1.7B
- SenseVoiceSmall
Scope and reproducibility
VoicePing publishes this evaluation of its own model. The results describe the dataset, model labels and scoring method reported above; they are not an independent certification or a guarantee for every language, recording or product workflow.
The article does not provide a complete public rerun package with all source data, exact service versions, requests, scoring code and environment details. Where those details are absent, do not infer them or treat a small score difference as statistically established. To evaluate your use case, test representative permitted samples, record the model versions and settings, and review errors as well as aggregate scores.


