To address listener fatigue caused by audiobook narration that relies on a small set of explicit, discrete emotion categories, we build a multi-modal continuous emotional space. Combined with a T5-based, contrastive-learning text-emotion-perception module, the approach significantly improves emotional similarity between the synthesized speech and its source text and surrounding context.