Text-aware and Context-aware Expressive Audiobook Speech Synthesis

Abstract

To address listener fatigue caused by audiobook narration that relies on a small set of explicit, discrete emotion categories, we build a multi-modal continuous emotional space. Combined with a T5-based, contrastive-learning text-emotion-perception module, the approach significantly improves emotional similarity between the synthesized speech and its source text and surrounding context.

Publication
Internship work at Huawei
Jiaxuan Liu
Jiaxuan Liu
M.Eng. Student in Information & Communication Engineering · Speech AI Researcher

My research focuses on text-to-speech, expressive & emotional speech synthesis, multimodal movie dubbing, and speech foundation models.