We select two samples generated by the baselines and one sample from the ESD dataset to compare with our sample.


Target : Target samples are provided from ESD dataset.

emospeech : Baseline emospeech model.

cosyvoice2 : Baseline cosyvoice2 model.

Our Emo-BPO : Our proposed Emo-BPO model.

Trulli
We propose Emotion Bidirectional Preference Optimization (Emo-BPO), a structured framework for diffusion-based emotional TTS. Emo-BPO constructs reordered same-text emotional pairs to jointly learn emotion-aligned and emotion-contrastive score functions. During inference, these branches are integrated through a contrastive guidance formulation compatible with CFG, strengthening target emotional trajectories while suppressing competing modes. Emo-BPO requires no additional annotations, reward models, or new training strategies and can be seamlessly integrated into existing diffusion-based TTS systems.


Sample 1 (Emotion: Angry)
Text: No, I burst the balloon!
Target emospeech cosyvoice2 our Emo-BPO
Samples


Sample 2 (Emotion: Surprise)
Text: The football teams give a tea party.
Target emospeech cosyvoice2 our Emo-BPO
Samples


Sample 3 (Emotion: Happy)
Text: That I owe my thanks to you.
Target emospeech cosyvoice2 our Emo-BPO
Samples


Sample 4 (Emotion: Neutral)
Text: Poor Tom now is dead.
Target emospeech cosyvoice2 our Emo-BPO
Samples


Sample 6 (Emotion: Sad)
Text: Must a name mean something?
Target emospeech cosyvoice2 our Emo-BPO
Samples