
In this article
Source Code:
- android-offline-transcribe — Android benchmark application and runtime implementations
- ios-mac-offline-transcribe — iOS and macOS benchmark application and runtime implementations
- windows-offline-transcribe — Windows benchmark application, CPU-only
Abstract
This speed-only benchmark compares speech-to-text model and engine configurations on four devices. The Android table has 16 configurations, iOS has 17, macOS has 15, and Windows has 12 entries, including failed or untimed entries where marked. Repeated model names represent different runtimes or quantizations, not additional distinct models. Android’s online speech recognizer is included as a control, not an offline result.
We measure inference speed, real-time factor (RTF), and runtime or memory limits. Transcription accuracy (WER/CER) is not evaluated. Moonshine Tiny has the shortest reported Android and Windows inference times; Parakeet TDT v3 with FluidAudio leads the reported iOS and macOS throughput. The tested Android Whisper Tiny configurations differ by about 51× in elapsed time. OOM-marked iOS values are partial measurements before a crash, not successful completed runs. All benchmark apps and results are open-source.
Motivation
Developers building voice-enabled edge applications face a combinatorial selection problem: dozens of ASR models (from 31 MB Whisper Tiny to 1.8 GB Qwen3 ASR), multiple inference engines (ONNX Runtime, CoreML, whisper.cpp, MLX), and 4+ target platforms — each combination producing vastly different speed and memory characteristics. Published model benchmarks typically report results on server GPUs, not the consumer mobile and laptop hardware where these models actually deploy.
This benchmark addresses the deployment choice problem directly: which model + engine combination delivers real-time transcription on each target platform, and what are the memory constraints? The results enable developers to select a model/engine pair based on measured data from their target device class, rather than extrapolating from GPU benchmarks.
Methodology
Android, iOS, and macOS benchmarks use the same 30-second WAV file containing a looped segment of the JFK inaugural address (16 kHz, mono, PCM 16-bit). Windows benchmarks use an 11-second excerpt from the same source (see Windows section for details).
How to read the measurements:
- Inference: Wall-clock time from the recognition call to its result. Model download and initial loading are excluded in the Android source; this is not the time from installing the app to receiving a transcript.
- tok/s: The source apps use this label for an output-words-per-second estimate, not a count from a shared tokenizer. A partial or incorrect transcript can change the numerator, so this value must be read with completion status.
- RTF: Real-Time Factor—the processing time divided by audio duration. For example, 12 seconds to process 30 seconds of audio gives RTF 0.4. This does not measure the delay before a live caption first appears.
- Size: Approximate model-package download size. It is not peak process memory, and different exports of the same model can have different sizes.
- PASS: The Android source checks for expected words from the JFK sample. Passing that check is not a word-error-rate evaluation and does not prove the whole transcript is correct.
The Android test notes also separate direct-file model runs from the built-in SpeechRecognizer loopback test. The latter plays audio through a speaker and captures it through a microphone on the tested API level, so its partial outputs are not a like-for-like speed comparison. No repeated-run distribution, battery measurement or transcription-accuracy score is reported here.
Devices:
| Device | Chip | RAM | OS |
|---|---|---|---|
| Samsung Galaxy S10 | Exynos 9820 | 8 GB | Android 12 (API 31) |
| iPad Pro 3rd gen | A12X Bionic | 4 GB | iOS 17+ |
| MacBook Air | Apple M4 | 32 GB | macOS 15+ |
| Laptop | Intel Core i5-1035G1 | 8 GB | Windows (CPU-only) |
Models and checkpoint variants
The images below show official model cards, distribution files, or API documentation captured on September 20, 2026. They identify the referenced implementation; download counts and publisher claims shown in them are not benchmark measurements. Open any image at full size or follow its source link.
The following 13 model entries group repeated runtime rows. Numbering continues into four operating-system speech interfaces below. It is an index, not a performance ranking; use the platform tables for exact configurations and completion status.
1. Moonshine Tiny

Moonshine Tiny is the smallest of the two original English-only Moonshine checkpoints, with 27 million parameters. Its role here is a lightweight starting point for local English transcription. The publisher’s model card targets devices with limited compute and memory; it does not establish support for Japanese or other languages.
In our sherpa-onnx configurations, Tiny had the shortest completed inference time on both Android and Windows: 1,363 ms for the 30-second Android clip and 435 ms for the 11-second Windows clip. On the Apple devices, it produced 37.3 output words/s on the iPad and 92.2 on the Mac. The approximately 125 MB download is a package size, not a measured RAM requirement.
For an English voice-command prototype, Tiny is a sensible baseline against which to measure the cost of larger models. Its speed leaves room for other application work, but this single speech sample cannot tell us whether it recognizes names, numbers or noisy commands correctly. Compare those errors with Moonshine Base before paying the larger model’s latency and storage cost.
2. Moonshine Base

Moonshine Base is the 61-million-parameter English checkpoint in the same family as Tiny. Keeping the language and runtime family similar makes this pair useful for examining the deployment cost of a larger model. The Moonshine model card distinguishes the two original sizes; neither entry here represents a newer multilingual or streaming Moonshine release.
On Android, Base took 2,251 ms versus Tiny’s 1,363 ms—about 1.65 times as long for the same clip. Its listed package is roughly 290 MB, compared with Tiny’s 125 MB. Windows showed a smaller timing gap, 534 ms versus 435 ms. Base also trailed Tiny in the reported Apple throughput: 31.3 versus 37.3 words/s on iOS, and 59.3 versus 92.2 on macOS.
These results quantify an extra runtime and download cost. They do not establish the benefit that might justify it, because this study did not score transcript accuracy. Base belongs in an English-only shortlist when a separate evaluation shows that its transcripts correct errors made by Tiny; parameter count alone is not evidence of that improvement.
3. Whisper Tiny

Whisper Tiny is OpenAI’s 39-million-parameter checkpoint. Unlike the original Moonshine entries, the multilingual Tiny checkpoint can be used for multiple languages. The Whisper model card distinguishes multilingual and English-only variants, so the exact checkpoint matters as much as the family name.
This is the clearest example of why a model name alone cannot predict application speed. The tested Android sherpa-onnx package completed in 2,068 ms, while the whisper.cpp path took 105,596 ms. The listed packages also differ—approximately 100 MB and 31 MB—so the 51× timing ratio describes these complete configurations, not a controlled experiment isolating only the inference library.
On the iPad, the whisper.cpp row reached 37.8 words/s, while the WhisperKit row reached 4.5. That reversal relative to the Android result is a reason to benchmark the actual platform integration. Tiny is useful as a compact multilingual baseline, but choose a checkpoint, quantization and runtime together. This study supplies speed observations, not a multilingual recognition-quality ranking.
4. Whisper Base

Whisper Base increases the model size to 74 million parameters. OpenAI publishes separate multilingual and English-only variants, described in the Whisper model card . Here, the Android row named “Whisper Base” is multilingual; the separate Base.en row is covered next. The Apple tables also label the WhisperKit Base configuration as English, so those rows should not be treated as interchangeable checkpoint packages.
The multilingual Android sherpa-onnx run completed in 4,038 ms, compared with Tiny’s 2,068 ms. On Windows, the Base whisper.cpp row took 6,501 ms for the 11-second clip and remained faster than the clip’s duration. On the 4 GB iPad, whisper.cpp completed at 13.8 words/s, but the tested WhisperKit Base configuration ran out of memory.
For deployment, that iPad result is more consequential than a headline throughput value: an incomplete run cannot serve a user. Base can be considered where Tiny’s transcripts are inadequate, but the target-device memory check must include the chosen runtime and the rest of the app. A 74M parameter count does not guarantee that every conversion will fit.
5. Whisper Base.en

Base.en is the English-only 74-million-parameter Whisper checkpoint, rather than a setting that adds English to the multilingual Base model. OpenAI lists the distinction in the Whisper model card . It is relevant when the application only accepts English and the team wants to compare a specialized checkpoint with the multilingual alternative.
The two Android sherpa-onnx rows are close in this run: 3,917 ms for Base.en and 4,038 ms for multilingual Base, with both packages listed at approximately 160 MB. The difference is only 121 ms. Without repeated runs or a variability estimate, it would be misleading to turn that gap into a dependable speed advantage.
The useful next comparison is transcript quality on the application’s English audio, while holding the runtime and audio pipeline constant. This article provides no WER measurement to resolve that choice. If the product also needs non-English speech, the Base.en timing does not justify replacing a multilingual checkpoint.
6. Whisper Small

Whisper Small is a 244-million-parameter multilingual checkpoint. In this comparison it illustrates a practical boundary: a model that completes quickly enough on one device/runtime combination may fall behind the incoming audio on another. Its larger checkpoint also raises the cost of loading and retaining the model alongside the rest of an application.
The Android sherpa-onnx row completed the 30-second clip in 12,329 ms, giving RTF 0.41. Windows whisper.cpp needed 21,260 ms for an 11-second clip, giving RTF 1.933. On the iPad, whisper.cpp completed at 3.9 words/s; the tested WhisperKit configuration instead ran out of memory. Its displayed 6.3 words/s was recorded before failure and is not a usable completed-run result.
This is a candidate to evaluate when smaller Whisper checkpoints miss important speech, but the benchmark does not demonstrate that accuracy improvement. For a continuously recording Windows application, the observed processing time already exceeds the clip duration before other application work is counted. For batch transcription with acceptable waiting time, the same constraint has a different consequence. See the checkpoint family documentation and the platform-specific rows before choosing an export.
7. Whisper Large v3 Turbo

The Turbo entries belong to the Whisper Large v3 Turbo family. The benchmark lists 809 million parameters, but distributes the model through several conversions, including sherpa-onnx, WhisperKit and whisper.cpp. “Compressed” in these tables is part of a package label; it does not establish that two packages have the same precision, memory use or runtime behavior.
Android’s sherpa-onnx configuration completed in 17,930 ms, or RTF 0.60. The Windows whisper.cpp run took 92,845 ms for 11 seconds of audio, or RTF 8.440. On the 4 GB iPad, both listed WhisperKit Turbo configurations failed with out-of-memory errors. The two whisper.cpp variants completed but were marked slower than real time. On the 32 GB Mac, the listed WhisperKit configurations completed, at 1.9 and 1.5 words/s.
Those observations make device and runtime selection decisive. The Android result cannot be transferred to the Windows laptop, and the iPad failure cannot be generalized to every Apple device. Before adopting a large checkpoint for local transcription, first establish an accuracy benefit over smaller alternatives, then measure peak process memory and sustained processing on the intended hardware. The Whisper runtime documentation is a starting point for checking supported conversions and hardware paths.
8. SenseVoice Small

SenseVoice Small covers Mandarin, Cantonese, English, Japanese and Korean. Its publisher documentation also describes emotion and audio-event recognition, but those additional tasks were not evaluated in this benchmark. The tested speech-to-text package uses sherpa-onnx and is listed at about 240 MB.
It was the second-fastest completed direct-file configuration by elapsed time in both the Android and Windows tables: 1,725 ms for the 30-second Android clip and 462 ms for the 11-second Windows clip. The Apple rows report 15.6 words/s on iOS and 27.4 on macOS. The measured audio was English, so these values do not demonstrate equivalent speed or accuracy for the other four languages.
SenseVoice is worth including in a shortlist that specifically needs its five-language coverage and a relatively compact local package. For example, a Japanese/English application has a language-fit reason to evaluate it that the original English-only Moonshine checkpoints do not provide. The selection should then turn on actual Japanese and English transcripts, especially names and mixed-language speech, rather than assuming that the English timing result establishes recognition quality.
9. Parakeet TDT v3

Parakeet TDT v3 is NVIDIA’s 600-million-parameter multilingual checkpoint. The official model card lists 25 European languages and automatic language detection. That coverage includes English, French, German and Spanish, but does not make this a Japanese, Korean or Mandarin recognizer.
The runtime pairing is central to the result. Android used sherpa-onnx and completed in 2,841 ms. The iOS and macOS rows used FluidAudio with CoreML and reported 181.8 and 171.6 output words/s, respectively—the highest displayed completed-run throughput in those two platform tables. The Android package is approximately 671 MB; the Apple package is listed at approximately 600 MB.
For an Apple application handling the supported languages, the FluidAudio result makes this a useful configuration to evaluate early. It does not establish that iOS is faster than macOS in general, or that the model will maintain the same throughput on long, noisy recordings. The Windows result below is v2, an English checkpoint: it must not be cited as a Windows measurement of v3. Language fit, package choice and successful completion come before comparing the displayed speeds.
10. Parakeet TDT v2

Parakeet TDT v2 is the English checkpoint used in the Windows comparison. Although v2 and v3 are both listed at 600 million parameters, they are distinct checkpoints with different language coverage. NVIDIA’s v2 model card describes English transcription with punctuation, capitalization and timestamp prediction; those capabilities should not be confused with v3’s expanded multilingual coverage.
The Windows sherpa-onnx configuration processed the 11-second sample in 1,239 ms, at RTF 0.113 and 17.8 words/s. Its approximately 660 MB package is substantially larger than the Moonshine Tiny package. The benchmark notes punctuated output, which is useful context for someone choosing a transcript-producing system, although punctuation quality was not scored systematically.
This result supports considering v2 for English file transcription on the tested CPU class. It does not prove a quality lead over the faster Moonshine and SenseVoice rows. An application that needs readable transcripts should compare punctuation, names and omissions directly; an application that needs European languages beyond English should evaluate v3 separately rather than transferring this Windows timing to it.
11. Zipformer Streaming 20M

The Zipformer entry uses an English streaming checkpoint through sherpa-onnx. Streaming means the recognizer accepts incoming chunks rather than requiring the application to wait for a whole recording. The Android implementation describes 100 ms chunks, while the published checkpoint package identifies the actual export used.
The measured Android run took 3,568 ms and the Windows run 1,775 ms. The Apple rows report 39.7 words/s on iOS and 77.4 on macOS. Packages are listed at about 73 MB on Android/Windows and 46 MB for the Apple INT8 configuration, so the rows should not be interpreted as a single identical binary tested everywhere.
Zipformer is particularly relevant when the interface must show words while someone is still speaking. However, the table measures processing throughput; it does not report time to the first partial result, how often partial text changes, or the delay before a phrase is finalized. Those interaction measurements determine whether streaming captions feel responsive. A fast file-processing number alone cannot answer that question.
12. Qwen3 ASR 0.6B

Qwen3-ASR-0.6B is the smaller of the two Qwen3-ASR checkpoints. The official model card describes 30 languages plus 22 Chinese dialects; “52” in some summaries combines those categories. This benchmark used English audio and does not verify recognition performance across that complete set.
The Android runtime difference is large: 15,881 ms with ONNX INT8 versus 338,261 ms with the pure C/NEON path. The first completed within the 30-second audio duration, while the second took more than five minutes. That is a comparison of these tested configurations, not evidence that every ONNX conversion is faster. On iOS, the two reported throughputs were close—5.4 and 5.6 words/s. On macOS, ONNX reached 8.0 versus 5.7 for the C path.
The macOS MLX package is listed but was not benchmarked. Its smaller download size cannot be turned into a GPU performance claim. Qwen’s language coverage may justify evaluating it where smaller English-only models are unsuitable, but the older Android device result makes runtime selection a first-order deployment decision. Include download, load time and application memory in that evaluation; the inference-only table excludes the first two.
13. Omnilingual 300M

The Omnilingual entry records the tested 300M CTC export, with an approximately 365 MB package and a broad advertised language inventory. The exact sherpa-onnx distribution matters: a failure in this export and integration is not a measurement of every model offered under a similar name.
Android returned output that failed the benchmark’s expected-keyword check after 44,035 ms. The macOS row is also marked as producing incorrect English output. Windows recorded 2,360 ms, but has no reported words-per-second value. None of those observations establishes a successful English transcription result that should enter the speed shortlist.
The source repository associates the wrong-language behavior with the tested CTC setup. Treat that as a reported integration limitation to investigate, not a general claim that broad-language recognition cannot work. If this model is required for a language absent from the smaller alternatives, first reproduce a correct transcript using the intended language and exact export. Only after that correctness check does timing become useful. Returning quickly with unrelated text is not a successful transcription.
Built-in speech systems and controls
These entries are system interfaces, not downloadable model checkpoints. Availability and offline behavior depend on the installed recognizer and configuration.
14. Android Speech (Offline)

Android SpeechRecognizer is an operating-system interface, not a downloadable checkpoint with a fixed model size. The installed recognition service and language resources determine what is available. Android provides an on-device availability check ; the existence of the API alone does not guarantee that a particular device has the required recognizer or language resources.
The reported row took 3,615 ms and produced 1.38 words/s, but there is an important measurement difference. On the tested Android 12/API 31 device, the benchmark played the WAV through the speaker while the recognizer listened through the microphone . The resulting transcript was partial and depended on the acoustic environment. Most other rows passed the audio directly to a local inference engine.
Consequently, this row should be kept separate from the direct-file speed ranking. Its practical appeal is using the device’s existing speech service without bundling a separate model. Before relying on that path for an offline feature, check the actual device’s availability and language resources, then verify recognition with the network disconnected. The table’s broad language-count label is not an offline-support guarantee.
15. Android Speech (Online)

The online SpeechRecognizer row is a network-dependent control, included to show what the Android system interface returned in the same test setup. It is not an offline alternative and should not be used to support a promise that an application works without connectivity.
It returned in 3,591 ms, close to the offline row’s 3,615 ms, with 1.39 reported words/s. That similarity does not demonstrate that online and offline recognition have the same latency. The benchmark’s Android notes explain that both system-recognizer rows used speaker-to-microphone loopback on this API level and returned partial transcripts. Their output and capture path differ from the direct-file model runs.
This control is useful when deciding whether a system service is sufficient for an application that permits network access. It offers little evidence about a downloadable model’s throughput or a fully offline deployment. A fair comparison would use the same audio input path, check transcript completeness, and account for the service and network conditions. Keep the current result as context rather than treating its near-identical timing as a model-level finding.
16. Apple Speech

Apple Speech exposes recognition through SFSpeechRecognizer rather than requiring the application to ship one of the model packages listed above. That distinction changes deployment work: an app integrates a system service and checks its capabilities, instead of selecting a specific checkpoint and quantization.
The macOS row reports 13.1 output words/s. It does not provide the full evidence needed to classify that run as strictly offline, and this article has no corresponding iOS Apple Speech timing. The “50+ languages” label is an inventory description, not a claim that every language is available for on-device recognition on every Mac or iPhone.
Apple documents a request property for requiring on-device recognition . A product with a strict local-processing requirement should verify support for its recognizer and locale before enabling that requirement, and then check the result without a network connection. The current row is useful as a system-service reference, but cannot settle that deployment requirement by itself. Nor does its throughput establish transcript quality relative to the downloadable models.
17. Windows Speech

Windows Speech is the operating-system speech interface included in the Windows application inventory. It has no reported inference time, words-per-second value or real-time factor in this benchmark. Those blank cells mean the entry was not timed; they must not be read as zero latency, a failed run, or a performance result comparable with the model rows.
The Windows benchmark repository identifies language support as dependent on installed packs. Microsoft’s SpeechRecognizer documentation describes the system API, which is a different deployment route from packaging Whisper or an ONNX model.
There is therefore no evidence here to recommend or reject Windows Speech on speed. A meaningful evaluation would first establish a working recognizer for the target language, feed it an equivalent sample, and record completion, transcript output and elapsed time. Until that exists, keep this entry in the availability inventory and exclude it from all measured rankings.
Android Results
Device: Samsung Galaxy S10, Android 12, API 31
| Model | Engine | Params | Size | Languages | Inference | tok/s | RTF | Result |
|---|---|---|---|---|---|---|---|---|
| Moonshine Tiny | sherpa-onnx | 27M | ~125 MB | English | 1,363 ms | 42.55 | 0.05 | ✅ |
| SenseVoice Small | sherpa-onnx | 234M | ~240 MB | zh/en/ja/ko/yue | 1,725 ms | 33.62 | 0.06 | ✅ |
| Whisper Tiny | sherpa-onnx | 39M | ~100 MB | 99 languages | 2,068 ms | 27.08 | 0.07 | ✅ |
| Moonshine Base | sherpa-onnx | 61M | ~290 MB | English | 2,251 ms | 25.77 | 0.08 | ✅ |
| Parakeet TDT 0.6B v3 | sherpa-onnx | 600M | ~671 MB | 25 European | 2,841 ms | 20.41 | 0.09 | ✅ |
| Android Speech (Offline) | SpeechRecognizer | System | Built-in | 50+ languages | 3,615 ms | 1.38 | 0.12 | ✅ |
| Android Speech (Online) | SpeechRecognizer | System | Built-in | 100+ languages | 3,591 ms | 1.39 | 0.12 | ✅ |
| Zipformer Streaming | sherpa-onnx streaming | 20M | ~73 MB | English | 3,568 ms | 16.26 | 0.12 | ✅ |
| Whisper Base (.en) | sherpa-onnx | 74M | ~160 MB | English | 3,917 ms | 14.81 | 0.13 | ✅ |
| Whisper Base | sherpa-onnx | 74M | ~160 MB | 99 languages | 4,038 ms | 14.36 | 0.13 | ✅ |
| Whisper Small | sherpa-onnx | 244M | ~490 MB | 99 languages | 12,329 ms | 4.70 | 0.41 | ✅ |
| Qwen3 ASR 0.6B (ONNX) | ONNX Runtime INT8 | 600M | ~1.9 GB | 30 languages | 15,881 ms | 3.65 | 0.53 | ✅ |
| Whisper Turbo | sherpa-onnx | 809M | ~1.0 GB | 99 languages | 17,930 ms | 3.23 | 0.60 | ✅ |
| Whisper Tiny (whisper.cpp) | whisper.cpp GGML | 39M | ~31 MB | 99 languages | 105,596 ms | 0.55 | 3.52 | ✅ |
| Qwen3 ASR 0.6B (CPU) | Pure C/NEON | 600M | ~1.8 GB | 30 languages | 338,261 ms | 0.17 | 11.28 | ✅ |
| Omnilingual 300M | sherpa-onnx | 300M | ~365 MB | 1,600+ languages | 44,035 ms | 0.05 | 1.47 | ❌ |
15 of 16 entries passed the source’s keyword check; no Android out-of-memory failure was reported. The two built-in SpeechRecognizer entries returned partial transcripts through acoustic loopback. Omnilingual failed the English check in this setup. These statuses describe test completion and expected-keyword presence, not measured transcription accuracy.
Android Engine Comparison: Same Model, Different Backends
The two tested Whisper Tiny configurations have a large timing difference:
| Backend | Inference | tok/s | Speedup |
|---|---|---|---|
| sherpa-onnx (ONNX) | 2,068 ms | 27.08 | 51x |
| whisper.cpp (GGML) | 105,596 ms | 0.55 | 1x (baseline) |
The elapsed-time ratio is approximately 51× for these two configurations. Their listed packages differ in size and format, and the report does not isolate every export, quantization and runtime setting. Use this result to motivate testing the actual integration; do not treat it as a universal speed advantage of one library over the other.
iOS Results
Device: iPad Pro 3rd gen, A12X Bionic, 4 GB RAM
| Model | Engine | Params | Size | Languages | tok/s | Status |
|---|---|---|---|---|---|---|
| Parakeet TDT v3 | FluidAudio (CoreML) | 600M | ~600 MB (CoreML) | 25 European | 181.8 | ✅ |
| Zipformer 20M | sherpa-onnx streaming | 20M | ~46 MB (INT8) | English | 39.7 | ✅ |
| Whisper Tiny | whisper.cpp | 39M | ~31 MB (GGML Q5_1) | 99 languages | 37.8 | ✅ |
| Moonshine Tiny | sherpa-onnx offline | 27M | ~125 MB (INT8) | English | 37.3 | ✅ |
| Moonshine Base | sherpa-onnx offline | 61M | ~280 MB (INT8) | English | 31.3 | ✅ |
| Whisper Base | WhisperKit (CoreML) | 74M | ~150 MB (CoreML) | English | 19.6 | ❌ OOM on 4 GB |
| SenseVoice Small | sherpa-onnx offline | 234M | ~240 MB (INT8) | zh/en/ja/ko/yue | 15.6 | ✅ |
| Whisper Base | whisper.cpp | 74M | ~57 MB (GGML Q5_1) | 99 languages | 13.8 | ✅ |
| Whisper Small | WhisperKit (CoreML) | 244M | ~500 MB (CoreML) | 99 languages | 6.3 | ❌ OOM on 4 GB |
| Qwen3 ASR 0.6B | Pure C (ARM NEON) | 600M | ~1.8 GB | 30 languages | 5.6 | ✅ |
| Qwen3 ASR 0.6B (ONNX) | ONNX Runtime (INT8) | 600M | ~1.6 GB (INT8) | 30 languages | 5.4 | ✅ |
| Whisper Tiny | WhisperKit (CoreML) | 39M | ~80 MB (CoreML) | 99 languages | 4.5 | ✅ |
| Whisper Small | whisper.cpp | 244M | ~181 MB (GGML Q5_1) | 99 languages | 3.9 | ✅ |
| Whisper Large v3 Turbo (compressed) | WhisperKit (CoreML) | 809M | ~1 GB (CoreML) | 99 languages | 1.9 | ❌ OOM on 4 GB |
| Whisper Large v3 Turbo | WhisperKit (CoreML) | 809M | ~600 MB (CoreML) | 99 languages | 1.4 | ❌ OOM on 4 GB |
| Whisper Large v3 Turbo | whisper.cpp | 809M | ~547 MB (GGML Q5_0) | 99 languages | 0.8 | ⚠️ RTF >1 |
| Whisper Large v3 Turbo (compressed) | whisper.cpp | 809M | ~834 MB (GGML Q8_0) | 99 languages | 0.8 | ⚠️ RTF >1 |
Memory limit in the tested iPad setup: The listed WhisperKit Base, Small and Turbo configurations ran out of memory on this 4 GB device. Their displayed throughput was observed before failure. The listed whisper.cpp alternatives completed, although the Turbo rows were slower than real time. This is evidence about these configurations on this iPad, not a guarantee about every 4 GB device or runtime version.
macOS Results
Device: MacBook Air M4, 32 GB RAM
| Model | Engine | Params | Size | Languages | tok/s | Status |
|---|---|---|---|---|---|---|
| Parakeet TDT v3 | FluidAudio (CoreML) | 600M | ~600 MB (CoreML) | 25 European | 171.6 | ✅ |
| Moonshine Tiny | sherpa-onnx offline | 27M | ~125 MB (INT8) | English | 92.2 | ✅ |
| Zipformer 20M | sherpa-onnx streaming | 20M | ~46 MB (INT8) | English | 77.4 | ✅ |
| Moonshine Base | sherpa-onnx offline | 61M | ~280 MB (INT8) | English | 59.3 | ✅ |
| SenseVoice Small | sherpa-onnx offline | 234M | ~240 MB (INT8) | zh/en/ja/ko/yue | 27.4 | ✅ |
| Whisper Tiny | WhisperKit (CoreML) | 39M | ~80 MB (CoreML) | 99 languages | 24.7 | ✅ |
| Whisper Base | WhisperKit (CoreML) | 74M | ~150 MB (CoreML) | English | 23.3 | ✅ |
| Apple Speech | SFSpeechRecognizer | System | Built-in | 50+ languages | 13.1 | ✅ |
| Whisper Small | WhisperKit (CoreML) | 244M | ~500 MB (CoreML) | 99 languages | 8.7 | ✅ |
| Qwen3 ASR 0.6B (ONNX) | ONNX Runtime (INT8) | 600M | ~1.6 GB (INT8) | 30 languages | 8.0 | ✅ |
| Qwen3 ASR 0.6B | Pure C (ARM NEON) | 600M | ~1.8 GB | 30 languages | 5.7 | ✅ |
| Whisper Large v3 Turbo | WhisperKit (CoreML) | 809M | ~600 MB (CoreML) | 99 languages | 1.9 | ✅ |
| Whisper Large v3 Turbo (compressed) | WhisperKit (CoreML) | 809M | ~1 GB (CoreML) | 99 languages | 1.5 | ✅ |
| Qwen3 ASR 0.6B (MLX) | MLX (Metal GPU) | 600M | ~400 MB (4-bit) | 30 languages | — | Not benchmarked |
| Omnilingual 300M | sherpa-onnx offline | 300M | ~365 MB (INT8) | 1,600+ languages | 0.03 | ❌ English broken |
No out-of-memory failure was reported for the completed Mac runs. The listed WhisperKit configurations completed on the 32 GB MacBook Air. That does not make every inventory row successful: Omnilingual is marked as incorrect English output, and the MLX Qwen row was not measured.
Windows Results
Device: Intel Core i5-1035G1 @ 1.00 GHz (4C/8T), 8 GB RAM, CPU-only
| Model | Engine | Params | Size | Inference | Words/s | RTF |
|---|---|---|---|---|---|---|
| Moonshine Tiny | sherpa-onnx offline | 27M | ~125 MB | 435 ms | 50.6 | 0.040 |
| SenseVoice Small | sherpa-onnx offline | 234M | ~240 MB | 462 ms | 47.6 | 0.042 |
| Moonshine Base | sherpa-onnx offline | 61M | ~290 MB | 534 ms | 41.2 | 0.049 |
| Parakeet TDT v2 | sherpa-onnx offline | 600M | ~660 MB | 1,239 ms | 17.8 | 0.113 |
| Zipformer 20M | sherpa-onnx streaming | 20M | ~73 MB | 1,775 ms | 12.4 | 0.161 |
| Whisper Tiny | whisper.cpp | 39M | ~80 MB | 2,325 ms | 9.5 | 0.211 |
| Omnilingual 300M | sherpa-onnx offline | 300M | ~365 MB | 2,360 ms | — | 0.215 |
| Whisper Base | whisper.cpp | 74M | ~150 MB | 6,501 ms | 3.4 | 0.591 |
| Qwen3 ASR 0.6B | qwen-asr (C) | 600M | ~1.9 GB | 13,359 ms | 1.6 | 1.214 |
| Whisper Small | whisper.cpp | 244M | ~500 MB | 21,260 ms | 1.0 | 1.933 |
| Whisper Large v3 Turbo | whisper.cpp | 809M | ~834 MB | 92,845 ms | 0.2 | 8.440 |
| Windows Speech | Windows Speech API | N/A | 0 MB | — | — | — |
Tested with an 11-second JFK inauguration audio excerpt (22 words). All models run CPU-only on x86_64 (i5-1035G1, 4C/8T) — no GPU acceleration. Parakeet TDT v2 is notable for combining fast inference (17.8 words/s) with full punctuation output.
Limitations
- Speed only: This benchmark measures inference speed, not transcription accuracy (WER/CER). Accuracy varies by model, language, and audio conditions — developers should evaluate accuracy separately for their target use case.
- Single audio sample: All platforms use a single JFK inaugural address recording. Results may differ with other audio characteristics (noise, accents, domain-specific vocabulary).
- Windows clip length: Windows uses an 11-second audio clip vs 30 seconds on other platforms, so cross-platform speed comparisons should account for this difference.
- iOS OOM values: The tok/s values shown for OOM-marked iOS models were measured before the crash and do not represent complete successful runs.
Further Research
- Add accuracy benchmarks: Pair speed metrics with WER/CER across multilingual datasets and noisy speech conditions.
- Expand audio conditions: Add long-form audio, overlapping speakers, and domain vocabulary (meetings, call-center, industrial) instead of one speech sample.
- Quantization sweep: Benchmark INT8/INT4 and mixed-precision variants across engines to map memory/speed/accuracy trade-offs.
- Extend the Qwen3-ASR evaluation: Add accuracy measurements and additional devices for the already-tested 0.6B checkpoint, and measure the currently untimed MLX path.
- Power and thermal profiling: Add battery drain and sustained-performance measurements for continuous on-device transcription workloads.
Conclusion
Several tested model/runtime configurations finish faster than real time, but the best choice depends on the target device and completion status. The Android Whisper Tiny comparison shows how much the tested engine configuration can matter: 2,068 ms with sherpa-onnx versus 105,596 ms with whisper.cpp. On Apple devices, the displayed Parakeet v3/FluidAudio path leads throughput, while larger WhisperKit configurations fail on the 4 GB iPad.
Use the result for the intended platform, exclude failed and untimed rows from successful-speed rankings, and treat Android online speech as a network-dependent control. Recognition quality and target-language suitability still need separate evaluation because this study did not measure WER or CER.
References
Our Repositories:
- android-offline-transcribe — Android benchmark app (Apache 2.0)
- ios-mac-offline-transcribe — iOS/macOS benchmark app (Apache 2.0)
- windows-offline-transcribe — Windows benchmark app (Apache 2.0)
Models:
- Moonshine Tiny/Base — Useful Sensors, English-only
- Whisper Tiny/Base/Small/Turbo — OpenAI, 99 languages
- SenseVoice Small — FunAudioLLM, 5 languages
- Parakeet TDT 0.6B v3 — NVIDIA NeMo, 25 European languages
- Qwen3 ASR 0.6B — Alibaba Qwen, 30 languages
- Zipformer Streaming — Next-gen Kaldi, English
- Omnilingual 300M — MMS, 1,600+ languages
Inference Engines:
- sherpa-onnx — Next-gen Kaldi ONNX Runtime
- whisper.cpp — C/C++ port of OpenAI Whisper
- WhisperKit — CoreML Whisper for Apple
- MLX — Apple Machine Learning framework


