Skip to main content
Speech Recognition On-Device AI Benchmark Whisper Moonshine Parakeet SenseVoice Qwen3 ASR Offline Transcription Android iOS macOS Windows

Offline Speech-to-Text Speed Compared Across Four Platforms

Akinori Nakajima - VoicePing
A sound waveform becomes digital data beside four device panels labeled Android, iOS, macOS and Windows. An illustration of the benchmark platforms, without performance scores.
In this article

Source Code:

Abstract

This speed-only benchmark compares speech-to-text model and engine configurations on four devices. The Android table has 16 configurations, iOS has 17, macOS has 15, and Windows has 12 entries, including failed or untimed entries where marked. Repeated model names represent different runtimes or quantizations, not additional distinct models. Android’s online speech recognizer is included as a control, not an offline result.

We measure inference speed, real-time factor (RTF), and runtime or memory limits. Transcription accuracy (WER/CER) is not evaluated. Moonshine Tiny has the shortest reported Android and Windows inference times; Parakeet TDT v3 with FluidAudio leads the reported iOS and macOS throughput. The tested Android Whisper Tiny configurations differ by about 51× in elapsed time. OOM-marked iOS values are partial measurements before a crash, not successful completed runs. All benchmark apps and results are open-source.

Motivation

Developers building voice-enabled edge applications face a combinatorial selection problem: dozens of ASR models (from 31 MB Whisper Tiny to 1.8 GB Qwen3 ASR), multiple inference engines (ONNX Runtime, CoreML, whisper.cpp, MLX), and 4+ target platforms — each combination producing vastly different speed and memory characteristics. Published model benchmarks typically report results on server GPUs, not the consumer mobile and laptop hardware where these models actually deploy.

This benchmark addresses the deployment choice problem directly: which model + engine combination delivers real-time transcription on each target platform, and what are the memory constraints? The results enable developers to select a model/engine pair based on measured data from their target device class, rather than extrapolating from GPU benchmarks.

Methodology

Android, iOS, and macOS benchmarks use the same 30-second WAV file containing a looped segment of the JFK inaugural address (16 kHz, mono, PCM 16-bit). Windows benchmarks use an 11-second excerpt from the same source (see Windows section for details).

How to read the measurements:

  • Inference: Wall-clock time from the recognition call to its result. Model download and initial loading are excluded in the Android source; this is not the time from installing the app to receiving a transcript.
  • tok/s: The source apps use this label for an output-words-per-second estimate, not a count from a shared tokenizer. A partial or incorrect transcript can change the numerator, so this value must be read with completion status.
  • RTF: Real-Time Factor—the processing time divided by audio duration. For example, 12 seconds to process 30 seconds of audio gives RTF 0.4. This does not measure the delay before a live caption first appears.
  • Size: Approximate model-package download size. It is not peak process memory, and different exports of the same model can have different sizes.
  • PASS: The Android source checks for expected words from the JFK sample. Passing that check is not a word-error-rate evaluation and does not prove the whole transcript is correct.

The Android test notes also separate direct-file model runs from the built-in SpeechRecognizer loopback test. The latter plays audio through a speaker and captures it through a microphone on the tested API level, so its partial outputs are not a like-for-like speed comparison. No repeated-run distribution, battery measurement or transcription-accuracy score is reported here.

Devices:

DeviceChipRAMOS
Samsung Galaxy S10Exynos 98208 GBAndroid 12 (API 31)
iPad Pro 3rd genA12X Bionic4 GBiOS 17+
MacBook AirApple M432 GBmacOS 15+
LaptopIntel Core i5-1035G18 GBWindows (CPU-only)

Models and checkpoint variants

The images below show official model cards, distribution files, or API documentation captured on September 20, 2026. They identify the referenced implementation; download counts and publisher claims shown in them are not benchmark measurements. Open any image at full size or follow its source link.

The following 13 model entries group repeated runtime rows. Numbering continues into four operating-system speech interfaces below. It is an index, not a performance ranking; use the platform tables for exact configurations and completion status.

1. Moonshine Tiny

Moonshine Tiny official checkpoint reference
The exact Tiny checkpoint, with its model card and package metadata. Checkpoint source

Moonshine Tiny is the smallest of the two original English-only Moonshine checkpoints, with 27 million parameters. Its role here is a lightweight starting point for local English transcription. The publisher’s model card targets devices with limited compute and memory; it does not establish support for Japanese or other languages.

In our sherpa-onnx configurations, Tiny had the shortest completed inference time on both Android and Windows: 1,363 ms for the 30-second Android clip and 435 ms for the 11-second Windows clip. On the Apple devices, it produced 37.3 output words/s on the iPad and 92.2 on the Mac. The approximately 125 MB download is a package size, not a measured RAM requirement.

For an English voice-command prototype, Tiny is a sensible baseline against which to measure the cost of larger models. Its speed leaves room for other application work, but this single speech sample cannot tell us whether it recognizes names, numbers or noisy commands correctly. Compare those errors with Moonshine Base before paying the larger model’s latency and storage cost.

2. Moonshine Base

Moonshine Base official checkpoint reference
The Base checkpoint is a separate download from Tiny; compare its own runtime and memory requirements. Checkpoint source

Moonshine Base is the 61-million-parameter English checkpoint in the same family as Tiny. Keeping the language and runtime family similar makes this pair useful for examining the deployment cost of a larger model. The Moonshine model card distinguishes the two original sizes; neither entry here represents a newer multilingual or streaming Moonshine release.

On Android, Base took 2,251 ms versus Tiny’s 1,363 ms—about 1.65 times as long for the same clip. Its listed package is roughly 290 MB, compared with Tiny’s 125 MB. Windows showed a smaller timing gap, 534 ms versus 435 ms. Base also trailed Tiny in the reported Apple throughput: 31.3 versus 37.3 words/s on iOS, and 59.3 versus 92.2 on macOS.

These results quantify an extra runtime and download cost. They do not establish the benefit that might justify it, because this study did not score transcript accuracy. Base belongs in an English-only shortlist when a separate evaluation shows that its transcripts correct errors made by Tiny; parameter count alone is not evidence of that improvement.

3. Whisper Tiny

Whisper Tiny official checkpoint reference
OpenAI’s multilingual Tiny checkpoint. The benchmark tests this model through several different inference engines. Checkpoint source

Whisper Tiny is OpenAI’s 39-million-parameter checkpoint. Unlike the original Moonshine entries, the multilingual Tiny checkpoint can be used for multiple languages. The Whisper model card distinguishes multilingual and English-only variants, so the exact checkpoint matters as much as the family name.

This is the clearest example of why a model name alone cannot predict application speed. The tested Android sherpa-onnx package completed in 2,068 ms, while the whisper.cpp path took 105,596 ms. The listed packages also differ—approximately 100 MB and 31 MB—so the 51× timing ratio describes these complete configurations, not a controlled experiment isolating only the inference library.

On the iPad, the whisper.cpp row reached 37.8 words/s, while the WhisperKit row reached 4.5. That reversal relative to the Android result is a reason to benchmark the actual platform integration. Tiny is useful as a compact multilingual baseline, but choose a checkpoint, quantization and runtime together. This study supplies speed observations, not a multilingual recognition-quality ranking.

4. Whisper Base

Whisper Base official checkpoint reference
The multilingual Base checkpoint; the English-only Base.en variant has its own section below. Checkpoint source

Whisper Base increases the model size to 74 million parameters. OpenAI publishes separate multilingual and English-only variants, described in the Whisper model card . Here, the Android row named “Whisper Base” is multilingual; the separate Base.en row is covered next. The Apple tables also label the WhisperKit Base configuration as English, so those rows should not be treated as interchangeable checkpoint packages.

The multilingual Android sherpa-onnx run completed in 4,038 ms, compared with Tiny’s 2,068 ms. On Windows, the Base whisper.cpp row took 6,501 ms for the 11-second clip and remained faster than the clip’s duration. On the 4 GB iPad, whisper.cpp completed at 13.8 words/s, but the tested WhisperKit Base configuration ran out of memory.

For deployment, that iPad result is more consequential than a headline throughput value: an incomplete run cannot serve a user. Base can be considered where Tiny’s transcripts are inadequate, but the target-device memory check must include the chosen runtime and the rest of the app. A 74M parameter count does not guarantee that every conversion will fit.

5. Whisper Base.en

Whisper Base.en official checkpoint reference
The English-only Base.en checkpoint. Its scope differs from the multilingual Base model. Checkpoint source

Base.en is the English-only 74-million-parameter Whisper checkpoint, rather than a setting that adds English to the multilingual Base model. OpenAI lists the distinction in the Whisper model card . It is relevant when the application only accepts English and the team wants to compare a specialized checkpoint with the multilingual alternative.

The two Android sherpa-onnx rows are close in this run: 3,917 ms for Base.en and 4,038 ms for multilingual Base, with both packages listed at approximately 160 MB. The difference is only 121 ms. Without repeated runs or a variability estimate, it would be misleading to turn that gap into a dependable speed advantage.

The useful next comparison is transcript quality on the application’s English audio, while holding the runtime and audio pipeline constant. This article provides no WER measurement to resolve that choice. If the product also needs non-English speech, the Base.en timing does not justify replacing a multilingual checkpoint.

6. Whisper Small

Whisper Small official checkpoint reference
The Small checkpoint’s model card. Package size alone does not determine peak runtime memory. Checkpoint source

Whisper Small is a 244-million-parameter multilingual checkpoint. In this comparison it illustrates a practical boundary: a model that completes quickly enough on one device/runtime combination may fall behind the incoming audio on another. Its larger checkpoint also raises the cost of loading and retaining the model alongside the rest of an application.

The Android sherpa-onnx row completed the 30-second clip in 12,329 ms, giving RTF 0.41. Windows whisper.cpp needed 21,260 ms for an 11-second clip, giving RTF 1.933. On the iPad, whisper.cpp completed at 3.9 words/s; the tested WhisperKit configuration instead ran out of memory. Its displayed 6.3 words/s was recorded before failure and is not a usable completed-run result.

This is a candidate to evaluate when smaller Whisper checkpoints miss important speech, but the benchmark does not demonstrate that accuracy improvement. For a continuously recording Windows application, the observed processing time already exceeds the clip duration before other application work is counted. For batch transcription with acceptable waiting time, the same constraint has a different consequence. See the checkpoint family documentation and the platform-specific rows before choosing an export.

7. Whisper Large v3 Turbo

Whisper Large v3 Turbo official checkpoint reference
The Large v3 Turbo checkpoint, identified separately from the original Large v3 model. Checkpoint source

The Turbo entries belong to the Whisper Large v3 Turbo family. The benchmark lists 809 million parameters, but distributes the model through several conversions, including sherpa-onnx, WhisperKit and whisper.cpp. “Compressed” in these tables is part of a package label; it does not establish that two packages have the same precision, memory use or runtime behavior.

Android’s sherpa-onnx configuration completed in 17,930 ms, or RTF 0.60. The Windows whisper.cpp run took 92,845 ms for 11 seconds of audio, or RTF 8.440. On the 4 GB iPad, both listed WhisperKit Turbo configurations failed with out-of-memory errors. The two whisper.cpp variants completed but were marked slower than real time. On the 32 GB Mac, the listed WhisperKit configurations completed, at 1.9 and 1.5 words/s.

Those observations make device and runtime selection decisive. The Android result cannot be transferred to the Windows laptop, and the iPad failure cannot be generalized to every Apple device. Before adopting a large checkpoint for local transcription, first establish an accuracy benefit over smaller alternatives, then measure peak process memory and sustained processing on the intended hardware. The Whisper runtime documentation is a starting point for checking supported conversions and hardware paths.

8. SenseVoice Small

SenseVoice Small official checkpoint reference
FunAudioLLM’s SenseVoiceSmall model card. The benchmark below measures the selected speech-recognition implementation. Checkpoint source

SenseVoice Small covers Mandarin, Cantonese, English, Japanese and Korean. Its publisher documentation also describes emotion and audio-event recognition, but those additional tasks were not evaluated in this benchmark. The tested speech-to-text package uses sherpa-onnx and is listed at about 240 MB.

It was the second-fastest completed direct-file configuration by elapsed time in both the Android and Windows tables: 1,725 ms for the 30-second Android clip and 462 ms for the 11-second Windows clip. The Apple rows report 15.6 words/s on iOS and 27.4 on macOS. The measured audio was English, so these values do not demonstrate equivalent speed or accuracy for the other four languages.

SenseVoice is worth including in a shortlist that specifically needs its five-language coverage and a relatively compact local package. For example, a Japanese/English application has a language-fit reason to evaluate it that the original English-only Moonshine checkpoints do not provide. The selection should then turn on actual Japanese and English transcripts, especially names and mixed-language speech, rather than assuming that the English timing result establishes recognition quality.

9. Parakeet TDT v3

Parakeet TDT v3 official checkpoint reference
NVIDIA’s multilingual v3 checkpoint. FluidAudio and sherpa-onnx results remain separate runtime configurations. Checkpoint source

Parakeet TDT v3 is NVIDIA’s 600-million-parameter multilingual checkpoint. The official model card lists 25 European languages and automatic language detection. That coverage includes English, French, German and Spanish, but does not make this a Japanese, Korean or Mandarin recognizer.

The runtime pairing is central to the result. Android used sherpa-onnx and completed in 2,841 ms. The iOS and macOS rows used FluidAudio with CoreML and reported 181.8 and 171.6 output words/s, respectively—the highest displayed completed-run throughput in those two platform tables. The Android package is approximately 671 MB; the Apple package is listed at approximately 600 MB.

For an Apple application handling the supported languages, the FluidAudio result makes this a useful configuration to evaluate early. It does not establish that iOS is faster than macOS in general, or that the model will maintain the same throughput on long, noisy recordings. The Windows result below is v2, an English checkpoint: it must not be cited as a Windows measurement of v3. Language fit, package choice and successful completion come before comparing the displayed speeds.

10. Parakeet TDT v2

Parakeet TDT v2 official checkpoint reference
NVIDIA’s English v2 checkpoint. A notice about v3 does not make these two checkpoints interchangeable. Checkpoint source

Parakeet TDT v2 is the English checkpoint used in the Windows comparison. Although v2 and v3 are both listed at 600 million parameters, they are distinct checkpoints with different language coverage. NVIDIA’s v2 model card describes English transcription with punctuation, capitalization and timestamp prediction; those capabilities should not be confused with v3’s expanded multilingual coverage.

The Windows sherpa-onnx configuration processed the 11-second sample in 1,239 ms, at RTF 0.113 and 17.8 words/s. Its approximately 660 MB package is substantially larger than the Moonshine Tiny package. The benchmark notes punctuated output, which is useful context for someone choosing a transcript-producing system, although punctuation quality was not scored systematically.

This result supports considering v2 for English file transcription on the tested CPU class. It does not prove a quality lead over the faster Moonshine and SenseVoice rows. An application that needs readable transcripts should compare punctuation, names and omissions directly; an application that needs European languages beyond English should evaluate v3 separately rather than transferring this Windows timing to it.

11. Zipformer Streaming 20M

Zipformer Streaming 20M official checkpoint reference
The dated sherpa-onnx distribution links back to its original streaming model and training code. Checkpoint source

The Zipformer entry uses an English streaming checkpoint through sherpa-onnx. Streaming means the recognizer accepts incoming chunks rather than requiring the application to wait for a whole recording. The Android implementation describes 100 ms chunks, while the published checkpoint package identifies the actual export used.

The measured Android run took 3,568 ms and the Windows run 1,775 ms. The Apple rows report 39.7 words/s on iOS and 77.4 on macOS. Packages are listed at about 73 MB on Android/Windows and 46 MB for the Apple INT8 configuration, so the rows should not be interpreted as a single identical binary tested everywhere.

Zipformer is particularly relevant when the interface must show words while someone is still speaking. However, the table measures processing throughput; it does not report time to the first partial result, how often partial text changes, or the delay before a phrase is finalized. Those interaction measurements determine whether streaming captions feel responsive. A fast file-processing number alone cannot answer that question.

12. Qwen3 ASR 0.6B

Qwen3 ASR 0.6B official checkpoint reference
The official 0.6B checkpoint page. Its family overview also discusses 1.7B; only the listed 0.6B configuration belongs to this benchmark. Checkpoint source

Qwen3-ASR-0.6B is the smaller of the two Qwen3-ASR checkpoints. The official model card describes 30 languages plus 22 Chinese dialects; “52” in some summaries combines those categories. This benchmark used English audio and does not verify recognition performance across that complete set.

The Android runtime difference is large: 15,881 ms with ONNX INT8 versus 338,261 ms with the pure C/NEON path. The first completed within the 30-second audio duration, while the second took more than five minutes. That is a comparison of these tested configurations, not evidence that every ONNX conversion is faster. On iOS, the two reported throughputs were close—5.4 and 5.6 words/s. On macOS, ONNX reached 8.0 versus 5.7 for the C path.

The macOS MLX package is listed but was not benchmarked. Its smaller download size cannot be turned into a GPU performance claim. Qwen’s language coverage may justify evaluating it where smaller English-only models are unsuitable, but the older Android device result makes runtime selection a first-order deployment decision. Include download, load time and application memory in that evaluation; the inference-only table excludes the first two.

13. Omnilingual 300M

Omnilingual 300M official checkpoint reference
The tested distribution’s file inventory identifies model.int8.onnx and its 365 MB download size. Checkpoint source

The Omnilingual entry records the tested 300M CTC export, with an approximately 365 MB package and a broad advertised language inventory. The exact sherpa-onnx distribution matters: a failure in this export and integration is not a measurement of every model offered under a similar name.

Android returned output that failed the benchmark’s expected-keyword check after 44,035 ms. The macOS row is also marked as producing incorrect English output. Windows recorded 2,360 ms, but has no reported words-per-second value. None of those observations establishes a successful English transcription result that should enter the speed shortlist.

The source repository associates the wrong-language behavior with the tested CTC setup. Treat that as a reported integration limitation to investigate, not a general claim that broad-language recognition cannot work. If this model is required for a language absent from the smaller alternatives, first reproduce a correct transcript using the intended language and exact export. Only after that correctness check does timing become useful. Returning quickly with unrelated text is not a successful transcription.

Built-in speech systems and controls

These entries are system interfaces, not downloadable model checkpoints. Availability and offline behavior depend on the installed recognizer and configuration.

14. Android Speech (Offline)

Android Speech (Offline) official API documentation
Android documents an explicit on-device availability check. Availability still depends on the device and installed recognition service. Official API documentation

Android SpeechRecognizer is an operating-system interface, not a downloadable checkpoint with a fixed model size. The installed recognition service and language resources determine what is available. Android provides an on-device availability check ; the existence of the API alone does not guarantee that a particular device has the required recognizer or language resources.

The reported row took 3,615 ms and produced 1.38 words/s, but there is an important measurement difference. On the tested Android 12/API 31 device, the benchmark played the WAV through the speaker while the recognizer listened through the microphone . The resulting transcript was partial and depended on the acoustic environment. Most other rows passed the audio directly to a local inference engine.

Consequently, this row should be kept separate from the direct-file speed ranking. Its practical appeal is using the device’s existing speech service without bundling a separate model. Before relying on that path for an offline feature, check the actual device’s availability and language resources, then verify recognition with the network disconnected. The table’s broad language-count label is not an offline-support guarantee.

15. Android Speech (Online)

Android Speech (Online) official API documentation
Android’s SpeechRecognizer overview describes possible remote processing. This benchmark row is a network-dependent control. Official API documentation

The online SpeechRecognizer row is a network-dependent control, included to show what the Android system interface returned in the same test setup. It is not an offline alternative and should not be used to support a promise that an application works without connectivity.

It returned in 3,591 ms, close to the offline row’s 3,615 ms, with 1.39 reported words/s. That similarity does not demonstrate that online and offline recognition have the same latency. The benchmark’s Android notes explain that both system-recognizer rows used speaker-to-microphone loopback on this API level and returned partial transcripts. Their output and capture path differ from the direct-file model runs.

This control is useful when deciding whether a system service is sufficient for an application that permits network access. It offers little evidence about a downloadable model’s throughput or a fully offline deployment. A fair comparison would use the same audio input path, check transcript completeness, and account for the service and network conditions. Keep the current result as context rather than treating its near-identical timing as a model-level finding.

16. Apple Speech

Apple Speech official API documentation
Apple’s on-device request property and support prerequisite. This API reference does not establish which mode the recorded run used. Official API documentation

Apple Speech exposes recognition through SFSpeechRecognizer rather than requiring the application to ship one of the model packages listed above. That distinction changes deployment work: an app integrates a system service and checks its capabilities, instead of selecting a specific checkpoint and quantization.

The macOS row reports 13.1 output words/s. It does not provide the full evidence needed to classify that run as strictly offline, and this article has no corresponding iOS Apple Speech timing. The “50+ languages” label is an inventory description, not a claim that every language is available for on-device recognition on every Mac or iPhone.

Apple documents a request property for requiring on-device recognition . A product with a strict local-processing requirement should verify support for its recognizer and locale before enabling that requirement, and then check the result without a network connection. The current row is useful as a system-service reference, but cannot settle that deployment requirement by itself. Nor does its throughput establish transcript quality relative to the downloadable models.

17. Windows Speech

Windows Speech official API documentation
Microsoft’s built-in SpeechRecognizer API, distinct from bundling a downloadable model with an application. Official API documentation

Windows Speech is the operating-system speech interface included in the Windows application inventory. It has no reported inference time, words-per-second value or real-time factor in this benchmark. Those blank cells mean the entry was not timed; they must not be read as zero latency, a failed run, or a performance result comparable with the model rows.

The Windows benchmark repository identifies language support as dependent on installed packs. Microsoft’s SpeechRecognizer documentation describes the system API, which is a different deployment route from packaging Whisper or an ONNX model.

There is therefore no evidence here to recommend or reject Windows Speech on speed. A meaningful evaluation would first establish a working recognizer for the target language, feed it an equivalent sample, and record completion, transcript output and elapsed time. Until that exists, keep this entry in the availability inventory and exclude it from all measured rankings.

Android Results

Device: Samsung Galaxy S10, Android 12, API 31

Android Inference Speed — Tokens per Second

ModelEngineParamsSizeLanguagesInferencetok/sRTFResult
Moonshine Tinysherpa-onnx27M~125 MBEnglish1,363 ms42.550.05✅
SenseVoice Smallsherpa-onnx234M~240 MBzh/en/ja/ko/yue1,725 ms33.620.06✅
Whisper Tinysherpa-onnx39M~100 MB99 languages2,068 ms27.080.07✅
Moonshine Basesherpa-onnx61M~290 MBEnglish2,251 ms25.770.08✅
Parakeet TDT 0.6B v3sherpa-onnx600M~671 MB25 European2,841 ms20.410.09✅
Android Speech (Offline)SpeechRecognizerSystemBuilt-in50+ languages3,615 ms1.380.12✅
Android Speech (Online)SpeechRecognizerSystemBuilt-in100+ languages3,591 ms1.390.12✅
Zipformer Streamingsherpa-onnx streaming20M~73 MBEnglish3,568 ms16.260.12✅
Whisper Base (.en)sherpa-onnx74M~160 MBEnglish3,917 ms14.810.13✅
Whisper Basesherpa-onnx74M~160 MB99 languages4,038 ms14.360.13✅
Whisper Smallsherpa-onnx244M~490 MB99 languages12,329 ms4.700.41✅
Qwen3 ASR 0.6B (ONNX)ONNX Runtime INT8600M~1.9 GB30 languages15,881 ms3.650.53✅
Whisper Turbosherpa-onnx809M~1.0 GB99 languages17,930 ms3.230.60✅
Whisper Tiny (whisper.cpp)whisper.cpp GGML39M~31 MB99 languages105,596 ms0.553.52✅
Qwen3 ASR 0.6B (CPU)Pure C/NEON600M~1.8 GB30 languages338,261 ms0.1711.28✅
Omnilingual 300Msherpa-onnx300M~365 MB1,600+ languages44,035 ms0.051.47❌

15 of 16 entries passed the source’s keyword check; no Android out-of-memory failure was reported. The two built-in SpeechRecognizer entries returned partial transcripts through acoustic loopback. Omnilingual failed the English check in this setup. These statuses describe test completion and expected-keyword presence, not measured transcription accuracy.

Android Engine Comparison: Same Model, Different Backends

The two tested Whisper Tiny configurations have a large timing difference:

BackendInferencetok/sSpeedup
sherpa-onnx (ONNX)2,068 ms27.0851x
whisper.cpp (GGML)105,596 ms0.551x (baseline)

The elapsed-time ratio is approximately 51× for these two configurations. Their listed packages differ in size and format, and the report does not isolate every export, quantization and runtime setting. Use this result to motivate testing the actual integration; do not treat it as a universal speed advantage of one library over the other.

iOS Results

Device: iPad Pro 3rd gen, A12X Bionic, 4 GB RAM

iOS Inference Speed — Tokens per Second

ModelEngineParamsSizeLanguagestok/sStatus
Parakeet TDT v3FluidAudio (CoreML)600M~600 MB (CoreML)25 European181.8✅
Zipformer 20Msherpa-onnx streaming20M~46 MB (INT8)English39.7✅
Whisper Tinywhisper.cpp39M~31 MB (GGML Q5_1)99 languages37.8✅
Moonshine Tinysherpa-onnx offline27M~125 MB (INT8)English37.3✅
Moonshine Basesherpa-onnx offline61M~280 MB (INT8)English31.3✅
Whisper BaseWhisperKit (CoreML)74M~150 MB (CoreML)English19.6❌ OOM on 4 GB
SenseVoice Smallsherpa-onnx offline234M~240 MB (INT8)zh/en/ja/ko/yue15.6✅
Whisper Basewhisper.cpp74M~57 MB (GGML Q5_1)99 languages13.8✅
Whisper SmallWhisperKit (CoreML)244M~500 MB (CoreML)99 languages6.3❌ OOM on 4 GB
Qwen3 ASR 0.6BPure C (ARM NEON)600M~1.8 GB30 languages5.6✅
Qwen3 ASR 0.6B (ONNX)ONNX Runtime (INT8)600M~1.6 GB (INT8)30 languages5.4✅
Whisper TinyWhisperKit (CoreML)39M~80 MB (CoreML)99 languages4.5✅
Whisper Smallwhisper.cpp244M~181 MB (GGML Q5_1)99 languages3.9✅
Whisper Large v3 Turbo (compressed)WhisperKit (CoreML)809M~1 GB (CoreML)99 languages1.9❌ OOM on 4 GB
Whisper Large v3 TurboWhisperKit (CoreML)809M~600 MB (CoreML)99 languages1.4❌ OOM on 4 GB
Whisper Large v3 Turbowhisper.cpp809M~547 MB (GGML Q5_0)99 languages0.8⚠️ RTF >1
Whisper Large v3 Turbo (compressed)whisper.cpp809M~834 MB (GGML Q8_0)99 languages0.8⚠️ RTF >1

Memory limit in the tested iPad setup: The listed WhisperKit Base, Small and Turbo configurations ran out of memory on this 4 GB device. Their displayed throughput was observed before failure. The listed whisper.cpp alternatives completed, although the Turbo rows were slower than real time. This is evidence about these configurations on this iPad, not a guarantee about every 4 GB device or runtime version.

macOS Results

Device: MacBook Air M4, 32 GB RAM

macOS Inference Speed — Tokens per Second

ModelEngineParamsSizeLanguagestok/sStatus
Parakeet TDT v3FluidAudio (CoreML)600M~600 MB (CoreML)25 European171.6✅
Moonshine Tinysherpa-onnx offline27M~125 MB (INT8)English92.2✅
Zipformer 20Msherpa-onnx streaming20M~46 MB (INT8)English77.4✅
Moonshine Basesherpa-onnx offline61M~280 MB (INT8)English59.3✅
SenseVoice Smallsherpa-onnx offline234M~240 MB (INT8)zh/en/ja/ko/yue27.4✅
Whisper TinyWhisperKit (CoreML)39M~80 MB (CoreML)99 languages24.7✅
Whisper BaseWhisperKit (CoreML)74M~150 MB (CoreML)English23.3✅
Apple SpeechSFSpeechRecognizerSystemBuilt-in50+ languages13.1✅
Whisper SmallWhisperKit (CoreML)244M~500 MB (CoreML)99 languages8.7✅
Qwen3 ASR 0.6B (ONNX)ONNX Runtime (INT8)600M~1.6 GB (INT8)30 languages8.0✅
Qwen3 ASR 0.6BPure C (ARM NEON)600M~1.8 GB30 languages5.7✅
Whisper Large v3 TurboWhisperKit (CoreML)809M~600 MB (CoreML)99 languages1.9✅
Whisper Large v3 Turbo (compressed)WhisperKit (CoreML)809M~1 GB (CoreML)99 languages1.5✅
Qwen3 ASR 0.6B (MLX)MLX (Metal GPU)600M~400 MB (4-bit)30 languages—Not benchmarked
Omnilingual 300Msherpa-onnx offline300M~365 MB (INT8)1,600+ languages0.03❌ English broken

No out-of-memory failure was reported for the completed Mac runs. The listed WhisperKit configurations completed on the 32 GB MacBook Air. That does not make every inventory row successful: Omnilingual is marked as incorrect English output, and the MLX Qwen row was not measured.

Windows Results

Device: Intel Core i5-1035G1 @ 1.00 GHz (4C/8T), 8 GB RAM, CPU-only

Windows Inference Speed — Words per Second

ModelEngineParamsSizeInferenceWords/sRTF
Moonshine Tinysherpa-onnx offline27M~125 MB435 ms50.60.040
SenseVoice Smallsherpa-onnx offline234M~240 MB462 ms47.60.042
Moonshine Basesherpa-onnx offline61M~290 MB534 ms41.20.049
Parakeet TDT v2sherpa-onnx offline600M~660 MB1,239 ms17.80.113
Zipformer 20Msherpa-onnx streaming20M~73 MB1,775 ms12.40.161
Whisper Tinywhisper.cpp39M~80 MB2,325 ms9.50.211
Omnilingual 300Msherpa-onnx offline300M~365 MB2,360 ms—0.215
Whisper Basewhisper.cpp74M~150 MB6,501 ms3.40.591
Qwen3 ASR 0.6Bqwen-asr (C)600M~1.9 GB13,359 ms1.61.214
Whisper Smallwhisper.cpp244M~500 MB21,260 ms1.01.933
Whisper Large v3 Turbowhisper.cpp809M~834 MB92,845 ms0.28.440
Windows SpeechWindows Speech APIN/A0 MB———

Tested with an 11-second JFK inauguration audio excerpt (22 words). All models run CPU-only on x86_64 (i5-1035G1, 4C/8T) — no GPU acceleration. Parakeet TDT v2 is notable for combining fast inference (17.8 words/s) with full punctuation output.

Limitations

  • Speed only: This benchmark measures inference speed, not transcription accuracy (WER/CER). Accuracy varies by model, language, and audio conditions — developers should evaluate accuracy separately for their target use case.
  • Single audio sample: All platforms use a single JFK inaugural address recording. Results may differ with other audio characteristics (noise, accents, domain-specific vocabulary).
  • Windows clip length: Windows uses an 11-second audio clip vs 30 seconds on other platforms, so cross-platform speed comparisons should account for this difference.
  • iOS OOM values: The tok/s values shown for OOM-marked iOS models were measured before the crash and do not represent complete successful runs.

Further Research

  • Add accuracy benchmarks: Pair speed metrics with WER/CER across multilingual datasets and noisy speech conditions.
  • Expand audio conditions: Add long-form audio, overlapping speakers, and domain vocabulary (meetings, call-center, industrial) instead of one speech sample.
  • Quantization sweep: Benchmark INT8/INT4 and mixed-precision variants across engines to map memory/speed/accuracy trade-offs.
  • Extend the Qwen3-ASR evaluation: Add accuracy measurements and additional devices for the already-tested 0.6B checkpoint, and measure the currently untimed MLX path.
  • Power and thermal profiling: Add battery drain and sustained-performance measurements for continuous on-device transcription workloads.

Conclusion

Several tested model/runtime configurations finish faster than real time, but the best choice depends on the target device and completion status. The Android Whisper Tiny comparison shows how much the tested engine configuration can matter: 2,068 ms with sherpa-onnx versus 105,596 ms with whisper.cpp. On Apple devices, the displayed Parakeet v3/FluidAudio path leads throughput, while larger WhisperKit configurations fail on the 4 GB iPad.

Use the result for the intended platform, exclude failed and untimed rows from successful-speed rankings, and treat Android online speech as a network-dependent control. Recognition quality and target-language suitability still need separate evaluation because this study did not measure WER or CER.

References

Our Repositories:

Models:

Inference Engines:

Share this article

Try VoicePing for Free

Break language barriers with AI translation. Start with our free plan today.

Video
0:00 0:00