Skip to main content
Voice Translation Automatic Dubbing Audio Translation Video Translation

12 Voice Translation and AI Dubbing Tools Compared

Arun Kumar - VoicePing
A conference speakerphone and a studio recording microphone represent live translation and recorded dubbing.
In this article

Choose by the output you need: a conversation people can follow now, or a recording you can edit and publish later. Live translation tools prioritize continuous speech and captions. Dubbing products add translated audio to recorded media, with different levels of script editing, speaker control and lip sync.

This guide covers 12 services, grouped by those workflows. VoicePing publishes the guide; the numbers identify services, not a quality ranking. Official features and public prices were checked on September 18, 2026. The retained July audio files are historical examples, not a new controlled test.

Choose a workflow

Live speech translation

1. VoicePing — live meetings and listener captions

VoicePing’s official guide to translating meeting audio.
VoicePing’s official guide to translating meeting audio. Official source

VoicePing captures supported meeting, browser or microphone audio and supplies transcription and translation. Its guide covers Zoom, Microsoft Teams and Google Meet, listener links or QR codes, and playback of translated final utterances.

  • Useful for: A meeting or event where people need captions and translated speech while the discussion continues, without building an API application.
  • Check before choosing: Audio routing, speaking language and playback settings still need setup. Test the actual microphone and meeting application. Live playback does not produce an edited, lip-synced video master.

The official setup guide explains the audio controls and notes that the high-accuracy setting may be disabled during translated TTS auto-play. Choose the listening workflow before evaluating the result.

2. Gemini Live Translate — a streaming translation API

Google’s official Live Translate documentation.
Google’s official Live Translate documentation. Official source

Google documents gemini-3.5-live-translate-preview for continuous speech translation across 70+ languages. The guide links to an AI Studio trial interface and describes streamed translated audio with optional input and output transcripts.

  • Useful for: Exploring a browser demo, then building a dedicated live interpreter into a custom application.
  • Check before choosing: This is a preview translation model with audio input. The translation configuration does not provide the tools and instructions of a general conversational agent. Production integration and operating costs remain the developer’s responsibility.

Google’s pricing bills input and output audio separately. A shorter input does not guarantee equally short translated output; calculate each side of the stream.

3. OpenAI Realtime Translate — translated audio and transcript updates

OpenAI’s official GPT-Realtime-Translate model page.
OpenAI’s official GPT-Realtime-Translate model page. Official source

gpt-realtime-translate uses a dedicated realtime translation endpoint. The documented outputs are translated audio and transcript updates while source audio is still arriving; billing is based on audio duration.

  • Useful for: Developers who need translated speech and incremental text inside their own live-audio experience.
  • Check before choosing: An API supplies a component of the workflow. Your application must handle audio capture, playback, user access and error states. The model page does not describe a dubbing timeline, lip-sync editor or finished-video export.

The official price schedule lists $0.034 per minute for this specific translation model. Do not substitute the token rates or benchmark results of a different Realtime model.

4. Azure Speech Translation — Microsoft-managed speech workflows

Microsoft’s overview of speech translation and Live Interpreter.
Microsoft’s overview of speech translation and Live Interpreter. Official source

Azure provides Speech SDK and CLI workflows for translating audio streams. The documentation distinguishes standard real-time speech translation from Live Interpreter. Standard translation can expose source transcription and translated text; final results can be synthesized into speech.

  • Useful for: Organizations building around Azure resources, regional deployment choices and existing Microsoft engineering operations.
  • Check before choosing: Select the specific translation mode and its supported languages. Live Interpreter currently documents target-language transcription without source-language transcription. Speech synthesis, additional targets and other services can change the bill.

The pricing page separates Standard translation, Live Interpreter and video translation. The $2.50/hour figure in the explanatory guide is an illustrative calculation, not a verified quote for every region or mode.

Dubbing recorded audio and video

5. ElevenLabs — automatic dubbing and legacy Studio editing

ElevenLabs’ official guide distinguishes v2 dubbing and the v1 Studio.
ElevenLabs’ official guide distinguishes v2 dubbing and the v1 Studio. Official source

ElevenLabs documents Automatic Dubbing with v2 Alpha and a separate v1 Dubbing Studio. The current v2 model supports API use. The Studio offers transcript editing, speaker reassignment and per-clip regeneration, but is in maintenance mode.

  • Useful for: Producing translated audio or video, with a deliberate choice between the current automatic model and the legacy detailed editor.
  • Check before choosing: v2 transcript editing and audio regeneration through the API are Enterprise features. Do not assume the v1 editor is available for a v2 job, or that dubbing automatically supplies lip sync.

The API schedule lists v2 at $2.20/minute, separately from v1 rates. Free-tier v2 dubs are watermarked; the guide does not offer a v2 watermark-for-discount switch.

6. HeyGen — video translation with lip-sync choices

HeyGen’s official video translation product page.
HeyGen’s official video translation product page. Official source

HeyGen combines video translation, translated voices, subtitles and lip sync. Its product workflow includes an Edit & Review step before finalizing the translated script. This is useful when the visible speaker and the translated soundtrack both matter.

  • Useful for: Presenter-led training, product explanations and other recordings that need a localized video output.
  • Check before choosing: Audio-only translation and lip-sync modes have different costs. Check the chosen engine, review controls and plan. A convincing face-to-voice match still needs an independent language review.

Web subscriptions and API purchases are separate. Check the official API pricing for the selected translation mode; custom avatar generation is a different charge.

7. Rask — recurring localization and review workflows

Rask’s official monthly plans and localization allowances.
Rask’s official monthly plans and localization allowances. Official source

Rask offers translation, a built-in editor and plans for recurring publishing. Its comparison covers script and subtitle controls, glossaries, team review and batch workflows, with access varying by tier.

  • Useful for: A creator or team localizing a continuing library of videos and coordinating corrections across languages.
  • Check before choosing: An included “minute” is a usage credit. Translating into multiple languages consumes more minutes, and lip sync consumes additional minutes. Review team size, API access and monthly versus annual allowances.

Rask’s billing explanation gives a practical example: one minute translated into one target language plus Enhanced Lip-sync uses four minutes of allowance—one for translation and three for lip sync. That enhanced rate is described as beta.

8. Deepdub — production-oriented dubbing and voice control

Deepdub’s official FAQ describes its production services.
Deepdub’s official FAQ describes its production services. Official source

Deepdub’s FAQ describes dubbing, voice referencing, accent controls and a voice library, alongside work by language and post-production specialists. Its offering therefore needs a production brief, rather than a comparison based only on a text-to-speech unit rate.

  • Useful for: A studio or enterprise assessing managed localization, character voices and the handoff to post-production.
  • Check before choosing: Obtain a quote for the actual language pair, duration, speaker count, review, revisions and output files. Confirm whether the proposed scope is software access, managed production or both.

The official FAQ describes the service approach, but does not establish a single public price for a complete dubbing job. Evaluate a representative approved sample before committing a catalog.

9. RWS / Papercup — managed localization with human review

RWS’s official announcement of its acquisition of Papercup’s technology.
RWS’s official announcement of its acquisition of Papercup’s technology. Official source

RWS announced the acquisition of Papercup’s intellectual property in June 2025. Its announcement describes combining voice synthesis and editorial tools with language specialists who review tone, pacing and accuracy. Treat Papercup in this shortlist as part of that RWS production discussion.

  • Useful for: Enterprises commissioning localized training, communications or media with a defined human review and delivery process.
  • Check before choosing: Request the current RWS scope and commercial proposal. Confirm language direction, reviewers, revisions, voice permissions, timing and final deliverables. A legacy Papercup pricing page does not establish a current self-service subscription.

The acquisition announcement is evidence of the change in ownership of the technology, not a public quote or an independent quality test.

10. Dubverse — dubbing, subtitles and creator editing

Dubverse’s official site separates dubbing, subtitles and text-to-speech.
Dubverse’s official site separates dubbing, subtitles and text-to-speech. Official source

Dubverse presents dubbing, subtitles and text-to-speech as distinct workflows. Its paid-plan descriptions include transcript customization, segment rework and subtitle export; voice cloning and other features differ by plan.

  • Useful for: Creators comparing an editor-based dubbing workflow with separate subtitle and narration tasks.
  • Check before choosing: Check the language and voice you need, then estimate the credit consumption of dubbing and revisions. A text-to-speech API price per character does not cover the complete video-localization job.

The public pricing page lists a two-day trial without a credit card and separate Pro and Supreme credit plans. Its displayed billing options and offers need confirmation at checkout; do not read an annual or half-yearly promotion as a monthly contract.

11. Perso Dubbing — creator plans with script and lip-dubbing tools

Perso’s official pricing page, with monthly billing selected.
Perso’s official pricing page, with monthly billing selected. Official source

Perso now publishes self-service Dubbing plans, alongside enterprise options. Its Starter plan lists script editing, a glossary, lip dubbing, downloads and a small voice library, with credit and project limits.

  • Useful for: Creators who want a bounded monthly plan for short localized videos and the ability to revise the script.
  • Check before choosing: Starter includes seven dubbing minutes but limits each video to three minutes. Creator increases the allowance and project length. Check lip-sync offers, credit usage, regeneration terms and export resolution for the actual plan.

The pricing page distinguishes the regular Starter price of $6.99/month from a $3.49 first-month offer for new subscribers. The Free option includes one generation at signup; that is not an unlimited recurring dubbing allowance.

12. YouTube automatic dubbing — translated tracks for a channel

YouTube’s official automatic dubbing eligibility and review guide.
YouTube’s official automatic dubbing eligibility and review guide. Official source

YouTube can generate translated audio tracks for eligible videos and channels. Creators can preview the result and manage publication in YouTube Studio. This keeps distribution and the additional audio tracks within the channel workflow.

  • Useful for: A creator whose goal is to offer more audio languages on YouTube without maintaining a separate publishing pipeline.
  • Check before choosing: Available language directions and video eligibility vary. Videos over 120 minutes, unsupported source languages or unsuitable speech may be ineligible. The guide says automatic dubs cannot be directly edited.

Set the original language correctly and review the dub before publication . YouTube explicitly notes that accents, names, idioms, noise and recognition errors can affect the output.

Compare prices and billing units

Prices checked September 18, 2026. Dollar amounts are USD. Tax treatment is stated where the source specifies it; otherwise verify regional taxes and the final checkout amount. Subscription allowances, API rates and managed-production quotes cover different deliverables.

Public prices and the conditions that affect a complete job
Service / planPrice and billing unitAllowance and limits
1. VoicePing Free / Individual
Official terms
Free: $0. Individual: $31.50/month, billed monthly; tax included.90 online minutes/month on Free; 450 on Individual, one active account. No Individual trial. Meeting translation, not a dubbing export allowance.
2. Gemini Live Translate Preview
Official terms
Audio: about $0.0053/input minute plus $0.0315/output minute.Free tier available with limits. About $0.0368 for one input minute plus one output minute; billed by audio tokens, not a fixed finished-video price.
3. OpenAI Realtime Translate
Official terms
$0.034 per realtime audio minute.API usage charge for this translation model. Application hosting, development and media editing are separate.
4. Azure Speech Translation
Official terms
F0 Standard translation: 5 audio hours/month free. Paid rates depend on region and configuration.Standard translation bills by the second. Confirm synthesis and extra-language costs; Live Interpreter has separate input/output billing.
5. ElevenLabs Dubbing API
Official terms
v2: $2.20/minute. v1: $0.33 with watermark or $0.50 without. Taxes excluded.Version and subscription allowances matter. v2 Free output is watermarked; v2 API transcript editing/regeneration requires Enterprise.
6. HeyGen Video Translation API
Official terms
Audio-only Speed: $0.57/minute; Speed lip sync: $0.81; Precision lip sync: $1.50.API pay-as-you-go, separate from web subscriptions. Credits expire after 12 months; no recurring free API credits. Verify selected mode before purchase.
7. Rask Creator / Creator Pro
Official terms
Creator: $60/month for 25 minutes. Creator Pro: $150/month for 100; both billed monthly.Three-minute trial. Multiple target languages and lip sync consume additional allowance; annual offers use different prices and allowances.
8. Deepdub production services
Official terms
Project quote; no comparable complete-job rate established.Specify language pairs, duration, voice work, review, revisions and delivery. Do not substitute TTS API rates.
9. RWS / Papercup production
Official terms
Current project proposal required.Managed scope and human review must be agreed. A legacy Papercup page does not establish the current RWS contract.
10. Dubverse Pro / Supreme
Official terms
Displayed monthly: Pro $18, Supreme $30; each with 50 credits/month.Two-day trial, no credit card. Confirm dubbing-credit conversion, revisions, billing cycle and any promotional terms.
11. Perso Free / Starter / Creator
Official terms
Starter: $6.99/month; Creator: $29/month, monthly billing.Starter: 7 dubbing minutes, 3-minute video limit. Creator: 40 minutes, 15-minute limit. Free: one signup generation. First-month Starter offer: $3.49.
12. YouTube automatic dubbing
Official terms
No separate generation fee is listed in the reviewed help guide.Channel, video and language eligibility apply. Output is an additional YouTube audio track; this is not a general export-studio subscription.

Estimate the whole job: source length × target languages, then add the selected lip-sync mode, regeneration, human review and delivery work. With subscriptions, divide by the minutes you expect to use, not the maximum advertised allowance. A short trial can reveal workflow friction without proving production quality.

Review the retained audio evidence

The files below were retained with the original July 2026 article. Their durations and formats can be checked, but they were not produced under a shared, controlled comparison protocol. No new paid API run or bilingual listening evaluation is claimed in this update.

The earlier Gemini and OpenAI interface captures use different inputs and do not supply matched downloadable outputs. The earlier ElevenLabs images show its Speech to Text editor, so they do not establish Dubbing Studio behavior. Current service profiles above use official source captures instead.

Source — Korean, 60.0 seconds, mono 16 kHz PCM WAV

Download the retained Korean source .

Azure output — English, 68.46 seconds, en-US-JennyNeural

Download the retained Azure WAV . The original record identifies the voice as en-US-JennyNeural. The longer output illustrates why source duration and translated-audio duration can differ; it is not a latency measurement or a scored accuracy result.

Source used alongside the retained output — Korean, 60.0 seconds

Download the source file . Its presence beside the next file does not establish the settings used to generate that output.

Retained ElevenLabs output — 60.03 seconds, mono 44.1 kHz MP3

Download the retained MP3 . Its target language, model, voice and edit settings were not preserved. Listen only as an archived artifact; it cannot establish translation accuracy, voice preservation or a current-model comparison.

For a useful new evaluation, use the same approved source and target language for each service. Record the model, plan, settings, edits, output duration and full job cost. Have a fluent reviewer assess meaning, names, numbers, terminology and timing; assess lip sync separately. Record revision time as well as generation time.

FAQ

Can a live translation API replace a video dubbing editor?

A live API returns translated speech or text as audio arrives. A finished video also needs timing, speaker review, soundtrack handling, export and sometimes lip sync. Check which of those steps the chosen product supplies.

Does lip sync prove that the translation is correct?

No. Mouth movement and translation accuracy are different checks. Have a fluent reviewer assess meaning, names, numbers, terminology and tone before approving the final video.

Are all prices per minute directly comparable?

No. A vendor may charge for source audio, generated audio, each target language, editing credits or a subscription allowance. Include lip sync, regeneration, review and export when estimating the complete job.

Are the retained audio samples a controlled quality benchmark?

No. They are archived files with different settings and incomplete provenance. The ElevenLabs file does not establish its target language or model. Use a shared source, recorded settings and bilingual review for a new comparison.

Build a shortlist around the delivery

For a live meeting, evaluate the listener’s experience and audio setup. For an application, evaluate the API contract and operating cost. For a recording, evaluate the editable script, speaker handling, soundtrack and export. For managed production, agree on who reviews the language and who approves the final delivery.

Before publishing a dub, confirm permission to use the source and any reproduced voice, then approve the language and performance. A useful shortlist makes those responsibilities clear before the first full production job.

Share this article

Translate the conversation as it happens

Explore VoicePing for live meeting translation, captions and translated speech playback.

Video
0:00 0:00