Recent LLM-based TTS models have achieved striking gains in expressive speech synthesis, yet they still struggle to produce fine-grained, interpretable emotional speech. Conventional methods rely on discrete emotion labels to control category and intensity, which cannot capture the complexity and continuity of human emotional perception. Moreover, the lack of large-scale emotional speech datasets with balanced distributions and fine-grained emotional annotations often causes overfitting and impedes effective emotion control.
To address these issues, we propose UDDETTS, a universal LLM framework that unifies discrete and dimensional emotions for controllable emotional TTS. UDDETTS introduces the interpretable Arousal–Dominance–Valence (ADV) space and supports control via either discrete emotion labels or non-linearly quantized ADV tokens, enabling fine-grained, decoupled emotion control beyond traditional label-based methods. The framework comprises a neural codec language model, an OT-CFM module with an emotional mixture encoder, and a vocoder, augmented with an ADV predictor that enables end-to-end synthesis even when explicit annotations are unavailable. A semi-supervised training strategy unifies spontaneous and elicited emotional speech datasets — fusing label-only and ADV-annotated data — to mitigate emotional imbalance and the low controllable coverage of the ADV space. Across label-controlled, ADV-controlled, and end-to-end emotional TTS, UDDETTS achieves linear emotion control along three interpretable dimensions and significantly outperforms strong LLM-based TTS baselines.