Speech Synthesis · Diffusion Models · Language Models

VALL-E vs NaturalSpeech 2 vs SwanVoice: Three Bets on Zero-Shot TTS

VALL-E turns speech into codec-token language modeling with a 3-second prompt, NaturalSpeech 2 diffuses continuous codec latents for prosody and singing, and SwanVoice generates a whole 1-4 speaker dialogue in one pass.

VALL-E vs NaturalSpeech 2 vs SwanVoice: Three Bets on Zero-Shot TTS

The three bets

All three papers behind this comparison attack the same problem: synthesize speech for a speaker the model has never seen, from a short prompt. They disagree on what the fundamental unit of that problem is. VALL-E bets on the interface: once speech is compressed into discrete codec tokens, text-to-speech becomes language modeling, and every scaling habit that worked for text LMs applies to voices. NaturalSpeech 2 bets on the signal: prosody is continuous acoustic variation, so it keeps the codec’s continuous latent vectors and generates them with a diffusion model instead of predicting tokens left to right. SwanVoice bets on the object: a dialogue is one continuous piece of audio, and the only way to keep timbre, timing, and mood consistent across speaker turns is to generate the whole conversation in a single pass.

These are not three versions of one model where the newest wins. Each paper changes a different layer of the stack (representation, generator, or generation unit), and the evidence below shows how little of the evaluation is actually shared between them.

Key numbers

DimensionVALL-ENaturalSpeech 2SwanVoice
Training data60K hours of English speech44K hours of speech and singingPaired corpus SwanData-Speech, size not stated on our page
Speech representationDiscrete codec tokensContinuous codec latent vectors25 Hz VAE latents
GeneratorAutoregressive language modelDiffusion modelFlow-matching Diffusion Transformer
Zero-shot prompt3-second acoustic clipSpeech promptZero-shot with speaker-turn conditioning
Unit generatedSingle utteranceSingle utterance, plus singingOne monologue or 1-4 speaker dialogue in one pass
Headline claimBetter naturalness and speaker similarity than the prior zero-shot TTS it was tested againstGains in prosody similarity, timbre similarity, robustness, and voice quality over previous systemsHighest richness and hierarchy scores on SwanBench-Speech among all evaluated open-source baselines
Named weak axisSkipped or repeated tokens, long-form drift, plus consent and impersonation riskDiffusion sampling latency; singing is sensitive to pitch and rhythmWord-level content accuracy: it can mispronounce or drop words

Read the table for what it does not contain. There is no row where all three systems are scored on the same benchmark under the same metric. VALL-E and NaturalSpeech 2 each report gains against baselines they selected, under their own protocols, on their own data. SwanVoice’s headline numbers come from SwanBench-Speech, a benchmark proposed by the same team, and its richness and hierarchy metrics measure expressiveness rather than whether the right words were spoken. Data scale is also confounded with everything else: 60K hours of English speech versus 44K hours of speech plus singing are different training regimes, so neither number explains a quality gap by itself. This page therefore never ranks the three 1-2-3. It maps each bet to the use case it fits and names what each paper itself flags as its weak axis.

What each paper is actually optimizing

VALL-E optimizes the interface. Its contribution is less a single MOS table than a reframing: tokenize speech with a neural audio codec, then train a conditional language model to emit codec tokens from text and a prompt clip. Pretraining on 60K hours of English speech and enrollment from a 3-second recording made zero-shot voice cloning feel like in-context learning, and the same interface preserves emotion and room acoustics from the prompt. The failure mode is the autoregressive one: token sequences can skip, repeat, or drift over long outputs, and a system that clones a voice from 3 seconds of audio carries obvious consent and impersonation risk that shapes its release posture.

NaturalSpeech 2 optimizes acoustic naturalness. It keeps the codec but keeps the latents continuous, then generates them with a diffusion model conditioned on text and a speech prompt. Duration and pitch predictors are also prompt-aware, so speaking style transfers from the reference clip rather than being re-inferred token by token. The paper targets both zero-shot speech and zero-shot singing from 44K hours of data, on the argument that diffusion handles continuous prosodic variation more directly than discrete prediction. The price is on the serving side: diffusion sampling raises latency and compute questions, and singing synthesis is unforgiving about pitch and rhythm errors.

SwanVoice optimizes cross-turn consistency. It identifies three failure modes of stitching turn-by-turn TTS into a conversation: timbre drift between turns of the same speaker, loss of conversational rhythm, and emotion resetting at every boundary. Its fix is architectural: a 25 Hz VAE keeps multi-minute audio tractable, raw-text conditioning with pause-aware symbols and pinyin substitution preserves pronunciation hints, and a flow-matching Diffusion Transformer receives the speaker-turn structure so it generates the dialogue layout as one continuous latent. Training runs as a curriculum from monologue to mixed data to real dialogue, followed by a DiffusionNFT post-training step with a phone-level reward and a speaker-similarity reward. The named weak axis is word-level content accuracy, and the paper is direct that expressiveness leads while getting every word right lags.

When to use which

  • Single-utterance voice cloning with the simplest scaling story. The codec-LM lineage that VALL-E defined is the most mature bet: representation and training recipes are well understood, and a short prompt is enough. If your workload is one speaker per clip and you can manage the safety surface, start here.
  • Expressive prosody or singing. NaturalSpeech 2 is the only system in this set whose recipe explicitly covers zero-shot singing, and its prompt-aware duration and pitch predictors exist for prosody transfer. Budget for diffusion sampling cost and validate pitch-critical content separately.
  • Multi-speaker dialogue. SwanVoice’s one-pass generation is the only design here that models cross-turn consistency directly rather than patching it after stitching. It handles 1 to 4 speakers. Reach for it when conversational delivery matters more than transcript-perfect fidelity, and measure word error rate on your own text before trusting it for narration.
  • Anything accuracy-critical, whatever the paper. All three papers report their own wins on their own evaluations. For a production pick, run all candidates on your own text and report intelligibility, speaker similarity, and failure cases with equal visibility.

Limits and open questions

The biggest limit of this comparison is that it cannot be a ranking. The three papers measure different things on different data: VALL-E reports naturalness and speaker similarity against its chosen baseline, NaturalSpeech 2 reports gains across prosody, timbre, stability, and voice quality on 44K hours that include singing, and SwanVoice reports self-proposed expressiveness metrics on a self-built benchmark. None of those numbers transfers across rows.

A second open question is latency, and it splits by generator. Autoregressive codec LMs pay per token and drift on long outputs; diffusion pays sampling steps at inference; SwanVoice’s single-pass design generates minutes of dialogue at once, which is exactly where its consistency advantage comes from, but the abstract leaves streaming behavior and memory open. If your application is interactive, this dimension may matter more than any table row.

Finally, voice cloning safety is not symmetric across the three. VALL-E’s 3-second enrollment is precisely what makes it the strongest impersonation risk of the set, and the same property in any of its successors deserves the same release discipline.

FAQ

Which of VALL-E, NaturalSpeech 2, and SwanVoice was trained on the most data?

On the numbers reported in the papers covered here, VALL-E’s 60K hours of English speech edges NaturalSpeech 2’s 44K hours of speech and singing. The comparison is weaker than it looks: the hours cover different content, languages, and collection regimes, and data scale is entangled with every other design choice, so it does not predict output quality on its own.

Can the three models be ranked on one benchmark or leaderboard?

No, and that is the main caveat of this comparison. Each paper evaluates against baselines and metrics it selected. SwanVoice’s headline results come from SwanBench-Speech, introduced by the same work, with richness and hierarchy metrics that score expressiveness rather than word accuracy. A shared-harness evaluation across all three does not exist in these papers.

Why does SwanVoice score highest on richness and hierarchy yet still mispronounce words?

Because those metrics measure different things. Richness and hierarchy capture expressive delivery and prosodic structure, where generating a whole dialogue in one pass helps. Word-level content accuracy depends on pronunciation modeling, which is SwanVoice’s named weak axis; its DiffusionNFT post-training adds a phone-level reward specifically to claw that back. Treat the two metric families as reporting different axes, not one trade-off curve.

Is NaturalSpeech 2 the only one of the three that can synthesize singing?

Among the three papers on this page, yes: zero-shot singing synthesis is an explicit target of NaturalSpeech 2, trained on speech and singing together. VALL-E targets English speech, and SwanVoice targets monologue and dialogue speech. If singing is a requirement, that alone filters the set.

Which method should I pick for a production voice-cloning feature?

Match the generation unit to the product first: codec LM for single-utterance cloning, diffusion latents for prosody or singing, one-pass generation for dialogue. Then evaluate every shortlisted candidate on your own text with word error rate, speaker similarity, and latency measured yourself, because none of the papers’ headline numbers was produced on your data.