
In this article
Abstract
We benchmarked five emotional text-to-speech systems for Japanese and Chinese across six target emotions: neutral, happy, sad, angry, fear, and disgust. The evaluation uses neutral prompts so the requested emotion must come from speech style, not from emotionally loaded text. Each model generated 120 samples, for a 600-WAV main benchmark corpus across the five completed systems.
The strongest balanced candidate is Qwen3-TTS CustomVoice 1.7B: it has the best pooled SenseVoice accuracy among models with trustworthy Japanese and Chinese text output, the lowest mean CER, the best anchor hit rate, and strong NISQA-TTS naturalness. CosyVoice 300M Instruct is the naturalness leader, but emotion recognition is weak, especially in Japanese. IndexTTS-2 reaches a high pooled SenseVoice score, but its Japanese CER is too high to treat that result as reliable Japanese TTS.
The most important pattern is language and emotion imbalance: Chinese SenseVoice emotion accuracy was higher for four of the five models; IndexTTS-2 was the exception, with unreliable Japanese text fidelity. No model received a matching SenseVoice label for fear or disgust in this run; that is a classifier result, not a completed listening judgment.
Models and evaluated controls
The five numbered entries identify the evaluated checkpoints. Each generated 120 samples across Japanese, Chinese and six target emotions. The scores below are automatic diagnostics; they do not replace human listening tests.
1. Qwen3-TTS CustomVoice 1.7B
Evaluated CustomVoice checkpoint . The run used the predefined speaker Ryan and a natural-language emotion instruction alongside the raw sentence.
CER: Japanese 8.6%, Chinese 9.7%. It offered the strongest balanced result across text fidelity and the emotion diagnostics in this study. Its Chinese SenseVoice emotion accuracy was 53.3%, while Japanese was 15.0%; even this model did not solve all six emotions reliably. The contrast between low Japanese CER and low Japanese emotion accuracy is important: recognizable text alone did not make the requested style recognizable to the classifier. Its Chinese anchor margin was positive, but its Japanese margin was negative. These are reasons to shortlist the model for listening tests, not proof that listeners will hear the intended emotion consistently.
Median wall time was 4.20 seconds, median RTF 1.58 and median peak VRAM 8.13 GB in the tested path. That RTF does not demonstrate faster-than-playback full synthesis in this run. The model card identifies Ryan as a native-English voice and recommends a speaker’s native language for best quality, while allowing supported cross-language synthesis. A Japanese or Chinese native-voice comparison would be useful, but was not performed here. Do not attribute the language gap to Ryan alone or assume changing the voice fixes it.
2. CosyVoice 300M Instruct
Evaluated 300M Instruct checkpoint . This is the checkpoint named in the experiment, distinct from CosyVoice2. The run used the built-in Japanese and Chinese male speakers with natural-language instructions.
CER: Japanese 43.9%, Chinese 11.1%. It led the automatic naturalness diagnostic with mean NISQA-TTS 4.267, but Japanese SenseVoice emotion accuracy was only 1.7%. Natural-sounding output did not imply reliable text or emotional control. The Japanese CER was also much higher than Qwen’s, so the attractive naturalness score should not dominate selection. In Chinese, the two emotion diagnostics diverged: SenseVoice accuracy was 36.7%, while anchor hit rate reached 72.0%. A classifier label and similarity to an anchor centroid answer different questions; their numerical percentages are not interchangeable measures of listener agreement.
CosyVoice had the fastest measured median wall time at 2.26 seconds, median RTF 0.85 and the lowest median peak VRAM at 3.96 GB. That makes this configuration useful for a resource-conscious listening comparison, provided text fidelity and target emotion are judged independently. Keep the result attached to the 300M Instruct checkpoint and named voices. It does not rank newer CosyVoice releases or establish first-audio latency, sustained serving capacity or the quality of a different instruction strategy.
3. Fish Audio S1-mini
Evaluated S1-mini checkpoint
. Emotion control used inline markers such as (joyful) or (sad), without a speaker or emotion reference WAV.
CER: Japanese 12.7%, Chinese 16.8%. SenseVoice emotion accuracy was 6.7% for Japanese and 16.7% for Chinese. The markers did not reliably move the generated speech to the requested emotion in this setup. Its Japanese text fidelity was much better than the IndexTTS-2 Japanese condition, yet its emotion scores remained low. The Chinese confusion examples are specific: all ten happy targets were classified as neutral, and most fear and disgust targets were also neutral. This shows what failed under the inline-marker protocol, without establishing that every supported Fish conditioning method would behave the same way.
The measured median wall time was 7.06 seconds and median RTF 3.47. Median peak VRAM was 13.05 GB, the largest in this comparison, despite the small 0.80 GB sampled process RSS. The server-backed adapter had no CPU coverage, so the process figure is not a complete host-memory or CPU budget. For a follow-up, record the full serving process and compare the exact conditioning interface used. Neither a newer Fish model nor a reference-audio workflow inherits this S1-mini result.
4. VoxCPM2
Official VoxCPM project . The main run wrapped the control instruction inline before the text and supplied no prompt or reference WAV.
CER: Japanese 18.6%, Chinese 4.4%. It had the lowest Chinese CER, but emotion accuracy was 6.7% for Japanese and 35.0% for Chinese. Good Chinese text fidelity should be read separately from emotional control and the slower measured synthesis path. The 4.4% Chinese CER makes VoxCPM2 a useful text-fidelity comparison, but Chinese anchor hit rate was only 36.0% and its anchor margin remained negative. The Chinese examples frequently collapsed to neutral, including all ten disgust targets. A model can therefore preserve the words while still failing the requested style in this automatic screen.
The main run’s median wall time was 28.44 seconds, median RTF 9.84 and median peak VRAM 12.79 GB. These measurements describe the chosen adapter and generation settings, not an inherent limit of the architecture. The official project describes voice design and reference-conditioned modes; this test used an inline instruction without a reference WAV. A comparison of those control paths and accelerated runtimes could change the outcome, but no such gains are measured here. Keep control-method compatibility separate from a blanket judgment of the model family.
5. IndexTTS-2
IndexTTS-2 paper
. The run used dataset-derived speaker prompts from JVNV for Japanese and CSEMOTIONS for Chinese, with text emotion conditioning through emo_text.
CER: Japanese 91.0%, Chinese 10.3%. Its Japanese SenseVoice emotion score of 43.3% cannot establish reliable Japanese TTS when the spoken text is this inaccurate. The Japanese condition remains an experimental comparison, not a supported-language or production-quality claim. Its Chinese CER was far lower than its Japanese CER, yet Chinese SenseVoice emotion accuracy was only 16.7% and the anchor margin was negative. The Japanese classifier sometimes returned a plausible emotion even when content fidelity was poor. That is why an emotion score should be screened against intelligibility before it becomes evidence for an expressive voice application.
Median wall time was 26.39 seconds, median RTF 6.97 and median peak VRAM 7.29 GB in this configuration. This model also received dataset-derived speaker prompts, unlike the reference-free Fish and VoxCPM2 conditions, so the control inputs were not identical. Japanese compatibility and the actual generated speech need investigation before changing the text pipeline; the current results do not isolate a tokenizer defect or prove that a fix will restore Japanese quality. Human review should score spoken content, emotion and speaker character separately, using the same permitted reference conditions across reruns.
Motivation
Emotional TTS is not just a naturalness problem. A model can sound fluent and pleasant while failing to express the requested style. For product use cases such as multilingual avatars, customer support voices, training simulations, or expressive speech translation, we need to know whether a TTS system can keep three things aligned at once:
- It says the intended Japanese or Chinese sentence.
- It sounds natural enough to listen to.
- It expresses the requested emotion rather than collapsing into neutral speech or a nearby emotion.
CLAP-style audio-text similarity is useful for broad retrieval, but it is too indirect for a six-label emotional TTS benchmark. This evaluation combines discrete emotion recognition, continuous emotion anchors, transcription correctness, naturalness predictors, runtime, and listening samples. The goal is not to declare a final production winner from automatic metrics alone; it is to screen models and identify which systems deserve human listening tests.
Evaluation Methodology
The benchmark uses a balanced generation grid across language, emotion, and prompt text:
The same sentence is reused across all six emotions. This keeps the task clean: if a Japanese sentence says “The meeting starts at 10 a.m.” or a Chinese sentence says “The documents are on the desk,” the model cannot rely on emotional text content. It must express the requested emotion through speech.
Prompt Set
Example Japanese prompts:
| ID | Sentence |
|---|---|
ja_001 | 会議は午前十時に始まります。 |
ja_002 | 資料は机の上に置いてあります。 |
ja_003 | 明日の予定を確認してください。 |
ja_004 | 電車は三番線から出発します。 |
ja_005 | 受付で名前を伝えてください。 |
Example Chinese prompts:
| ID | Sentence |
|---|---|
zh_001 | 会议将在上午十点开始。 |
zh_002 | 资料已经放在桌子上。 |
zh_003 | 请确认明天的日程安排。 |
zh_004 | 列车将从三号站台出发。 |
zh_005 | 请在前台告知您的姓名。 |
Emotion Controls
| Target emotion | Control text |
|---|---|
neutral | Speak in a clear, neutral, natural voice. |
happy | Speak in a happy, warm, bright voice. |
sad | Speak in a sad, soft, slow, gentle voice. |
angry | Speak in an angry, tense, forceful voice. |
fear | Speak in a fearful, tense, trembling voice. |
disgust | Speak in a disgusted, displeased, rejecting voice. |
Each model receives the same target label and text, but the actual control interface is model-specific:
| Model | Speaker/reference input used | Emotion control |
|---|---|---|
qwen3_tts_customvoice_1_7b | Predefined CustomVoice speaker Ryan. | Raw sentence plus natural-language control instruction. |
cosyvoice_300m_instruct | Named built-in speaker: Japanese 日语男, Chinese 中文男. | Raw sentence plus natural-language control instruction. |
fish_audio_s1_mini | No speaker or emotion reference WAV. | Inline marker such as (joyful), (sad), (angry), (scared), or (disgusted). |
voxcpm2 | No prompt/reference WAV in the main run. | Control instruction wrapped inline before the text. |
indextts-2 | Dataset-derived speaker prompt WAVs: JVNV for Japanese, CSEMOTIONS for Chinese. | Raw sentence plus text emotion conditioning through emo_text. |
Metrics
- SenseVoice emotion accuracy: primary automatic screen. SenseVoice predictions are mapped to the six benchmark labels;
surprisedandunknowncount as non-matches. - emotion2vec anchor hit and margin: secondary diagnostic using human emotional-speech anchor centroids from CSEMOTIONS for Chinese and JVNV for Japanese.
- CER: faster-whisper-large-v3 transcription against the original prompt text, used to verify that emotional expression did not break the spoken content.
- NISQA-TTS: primary naturalness diagnostic for synthesized speech.
- UTMOS: secondary quality diagnostic; useful as a warning signal, but harsher and more out-of-domain for Japanese/Chinese.
- RTF: real-time factor for synthesis speed.
Results
Resource Usage
Resource metrics come from metrics/generation_runs.csv for the 600 successful generated rows. They are operational diagnostics rather than strict hardware benchmarks: GPU, VRAM, wall time, and RTF are populated for all completed rows, while CPU is not captured for server-backed adapters that run outside the sampled process tree.
| Model | Median wall time | Median RTF | Median peak VRAM | GPU util | GPU power | CPU | Median peak RSS |
|---|---|---|---|---|---|---|---|
cosyvoice_300m_instruct | 2.26s | 0.85 | 3.96 GB | 30.3% avg / 39.0% peak | 145.0W avg / 155.6W peak | 127.8% peak; 100% coverage | 5.54 GB |
qwen3_tts_customvoice_1_7b | 4.20s | 1.58 | 8.13 GB | 22.9% avg / 25.0% peak | 126.3W avg / 127.1W peak | 138.1% peak; 100% coverage | 6.22 GB |
fish_audio_s1_mini | 7.06s | 3.47 | 13.05 GB | 25.3% avg / 69.0% peak | 150.4W avg / 183.7W peak | not captured; 0% coverage | 0.80 GB |
indextts-2 | 26.39s | 6.97 | 7.29 GB | 18.2% avg / 100.0% peak | 131.3W avg / 199.6W peak | not captured; 0% coverage | 7.69 GB |
voxcpm2 | 28.44s | 9.84 | 12.79 GB | 12.3% avg / 100.0% peak | 106.7W avg / 191.5W peak | not captured; 0% coverage | 10.65 GB |
CosyVoice is the fastest and lowest-VRAM model in this run, but it is not the strongest emotion-control candidate. Qwen3-TTS requires more VRAM than CosyVoice but remains much faster than IndexTTS-2 and VoxCPM2 while keeping the best balance of emotion recognition and text fidelity. Fish Audio has a small process RSS footprint, but its GPU memory footprint is the largest of the completed models.
JA/ZH Metrics Overview
This split table is the quickest way to compare Japanese and Chinese behavior across the three core automatic checks: SenseVoice emotion accuracy, CER text fidelity, and emotion2vec anchor alignment.
| Model | JA SenseVoice | ZH SenseVoice | JA CER | ZH CER | JA anchor hit | ZH anchor hit | JA anchor margin | ZH anchor margin |
|---|---|---|---|---|---|---|---|---|
qwen3_tts_customvoice_1_7b | 15.0% | 53.3% | 8.6% | 9.7% | 40.0% | 64.0% | -0.06645 | 0.04480 |
indextts-2 | 43.3% | 16.7% | 91.0% | 10.3% | 38.0% | 30.0% | -0.08293 | -0.04063 |
voxcpm2 | 6.7% | 35.0% | 18.6% | 4.4% | 40.0% | 36.0% | -0.04479 | -0.02693 |
cosyvoice_300m_instruct | 1.7% | 36.7% | 43.9% | 11.1% | 24.0% | 72.0% | -0.05481 | 0.03796 |
fish_audio_s1_mini | 6.7% | 16.7% | 12.7% | 16.8% | 20.0% | 24.0% | -0.08972 | -0.09542 |
Chinese is generally easier for the automatic emotion metrics, but CER and emotion accuracy do not always move together. Qwen3-TTS keeps CER low in both languages, while IndexTTS-2 has the highest Japanese SenseVoice score and also the worst Japanese CER.
Text Fidelity (CER)
For text fidelity, Qwen3-TTS is the most stable JA/ZH result: Japanese CER is 8.6% and Chinese CER is 9.7%. IndexTTS-2 is the warning case. Its pooled emotion score looks competitive, but its Japanese CER reaches 91.0%, so the generated Japanese text path is not reliable enough in this setup.
CER is calculated from an automatic retranscription of the generated audio. It can reflect synthesis errors, recognizer errors and differences in text normalization, so it is a diagnostic rather than a manual transcript audit. Listen to the original waveform when investigating high-error samples; the table alone does not identify which stage caused the mismatch.
Emotion Accuracy
SenseVoice
Chinese SenseVoice accuracy was higher for four of the five models, with IndexTTS-2 the exception. For Qwen3-TTS, Chinese SenseVoice accuracy is 53.3% while Japanese is 15.0%, even though CER is low in both languages. That suggests the issue is not just intelligibility; the emotional cues recognized by SenseVoice are much weaker or less aligned in Japanese.
No fear or disgust target received a matching SenseVoice label: recall is 0.0% for both emotions across all evaluated model/language pairs. This does not establish that listeners would also fail to recognize every example. These labels often collapse into sad, neutral, angry, or unknown.
Rows are target emotions and columns are SenseVoice predictions. Green boxes mark the ideal diagonal.
Compact failure-mode highlights:
| Case | What happened | Why it matters |
|---|---|---|
indextts-2 / ja | happy -> sad 4/10; fear -> sad 5/10; disgust -> angry 10/10. | Emotion labels may look plausible even when Japanese text quality is unreliable. |
qwen3_tts_customvoice_1_7b / zh | happy -> neutral 5/10; fear -> sad 9/10; disgust -> neutral 9/10. | Qwen is the balanced winner, but hard emotions still collapse. |
cosyvoice_300m_instruct / ja | happy -> unknown 10/10; fear -> unknown 9/10; disgust -> unknown 8/10. | Naturalness does not guarantee recognizable emotional control. |
fish_audio_s1_mini / zh | happy -> neutral 10/10; fear -> neutral 9/10; disgust -> neutral 8/10. | Inline emotion markers did not reliably shift the generated prosody. |
voxcpm2 / zh | happy -> neutral 7/10; fear -> neutral 6/10; disgust -> neutral 10/10. | Prompt-driven control often collapsed into neutral speech. |
emotion2vec Anchors
The anchor metric tells a similar story to SenseVoice: Chinese anchors are more favorable than Japanese anchors. A positive margin means the generated audio is closer to the target emotion centroid than to the nearest non-target centroid. Qwen3-TTS has a positive Chinese margin, while every Japanese margin is negative.
Unlike SenseVoice, the anchor diagnostic is a centroid-similarity check rather than a label classifier, so the useful visual is the hit/margin split rather than a confusion matrix.
The anchor set was incomplete: Japanese neutral and Chinese disgust anchors were missing. Its coverage therefore differs from the full six-label SenseVoice screen. A high anchor hit rate cannot be read as success across all six emotions, and a negative mean margin describes relative similarity under these available anchors rather than a direct human judgment of expression.
Naturalness
| Model | Mean NISQA-TTS | Low NISQA-TTS <3.0 | Mean UTMOS | Low UTMOS <3.0 |
|---|---|---|---|---|
cosyvoice_300m_instruct | 4.267 | 0.0% | 3.282 | 20.8% |
indextts-2 | 4.063 | 11.7% | 2.078 | 93.3% |
qwen3_tts_customvoice_1_7b | 4.007 | 0.8% | 2.939 | 51.7% |
fish_audio_s1_mini | 3.935 | 3.3% | 2.932 | 55.8% |
voxcpm2 | 3.788 | 8.3% | 2.596 | 76.7% |
Naturalness and emotional correctness are different questions. CosyVoice is the clearest naturalness winner, but it is not the emotion-control winner. Qwen3-TTS is slightly behind CosyVoice on NISQA-TTS, but substantially better on the balanced emotion/intelligibility trade-off.
Both naturalness columns are model predictions, not mean ratings collected from listeners. Their disagreement is useful: IndexTTS-2 had mean NISQA-TTS 4.063 but mean UTMOS 2.078. The article already notes the latter predictor’s domain limitations for Japanese and Chinese. Neither score should override the text-fidelity check or be labeled a completed listening study.
Listening Examples
The table below uses the same prompt index for happy and angry in Japanese and Chinese. These clips are not a human listening test; they are qualitative anchors for the automatic metrics.
| Model | Language | Target | SenseVoice prediction | Sample |
|---|---|---|---|---|
qwen3_tts_customvoice_1_7b | JA | happy | unknown | |
qwen3_tts_customvoice_1_7b | JA | angry | angry | |
qwen3_tts_customvoice_1_7b | ZH | happy | neutral | |
qwen3_tts_customvoice_1_7b | ZH | angry | angry | |
cosyvoice_300m_instruct | JA | happy | unknown | |
cosyvoice_300m_instruct | JA | angry | unknown | |
cosyvoice_300m_instruct | ZH | happy | happy | |
cosyvoice_300m_instruct | ZH | angry | neutral | |
indextts-2 | JA | happy | sad | |
indextts-2 | JA | angry | surprised | |
indextts-2 | ZH | happy | neutral | |
indextts-2 | ZH | angry | neutral | |
fish_audio_s1_mini | JA | happy | happy | |
fish_audio_s1_mini | JA | angry | happy | |
fish_audio_s1_mini | ZH | happy | neutral | |
fish_audio_s1_mini | ZH | angry | neutral | |
voxcpm2 | JA | happy | unknown | |
voxcpm2 | JA | angry | angry | |
voxcpm2 | ZH | happy | happy | |
voxcpm2 | ZH | angry | angry |
Limitations
- Automatic emotion labels are not human judgment. SenseVoice is useful because it supports Japanese and Chinese and emits labels that map to the benchmark, but it can have classifier bias and language imbalance.
- Anchor metrics depend on the anchor datasets. Japanese anchors come from JVNV and Chinese anchors from CSEMOTIONS;
ja/neutralandzh/disgustanchors were missing in this run. - IndexTTS-2 Japanese is diagnostic, not production evidence. Its pooled emotion score looks strong, but Japanese CER is too high in this setup.
Further Research
- Run a small native-listener MOS/CMOS test for Qwen3-TTS and CosyVoice, with separate ratings for naturalness, emotion correctness, and text intelligibility.
- Investigate IndexTTS-2 Japanese compatibility, input handling and generated audio before a rerun. The current CER does not isolate a tokenizer defect; evaluate any proposed change instead of assuming it fixes the result.
- Add or curate missing
ja/neutralandzh/disgustemotion anchors. - Run a focused Chinese human check for
sad,angry,fear, anddisgust, where automatic metrics show strong differences between easy and hard labels. - Keep SenseVoice as an automatic screening metric, but make final production decisions with human listening tests.
Conclusion
For Japanese and Chinese emotional TTS, Qwen3-TTS CustomVoice 1.7B is the strongest balanced model in this benchmark. It does not solve every emotion, but it combines the best practical mix of emotion recognition, low CER, anchor hit rate, naturalness, and runtime.
CosyVoice 300M Instruct is the naturalness leader and remains worth testing in human listening studies, but it should not be treated as solved six-emotion control. IndexTTS-2 is diagnostically interesting, especially for Chinese, but the Japanese condition does not support reliable speech generation and needs investigation before a new evaluation.
The biggest open problem is not raw naturalness. It is reliable, language-consistent emotion control. Most models scored higher in Chinese than Japanese on the automatic emotion screen. The zero classifier recall for fear and disgust needs native-listener evaluation before it can support a conclusion about perceived emotion.


