Fun-CineForge: A Unified Dataset Pipeline and MLLM-based Model for Zero-Shot Movie Dubbing in Diverse Cinematic Scenes

Abstract

Movie dubbing synthesizes speech from scripts conditioned on video scenes, requiring accurate lip sync, faithful timbre transfer, and proper modeling of character identity and emotion. Existing approaches face two major limitations: (1) high-quality multimodal dubbing datasets are limited in scale, suffer from high word-error rates, sparse annotations, costly manual labeling, and are restricted to monologue scenes; (2) existing dubbing models rely solely on the lip region for audio–visual alignment, limiting them in complex live-action scenes and yielding suboptimal lip sync, speech quality, and emotional expressiveness.

We propose Fun-CineForge, comprising an end-to-end production pipeline for large-scale dubbing datasets and an MLLM-based dubbing model designed for diverse cinematic scenes. Using the pipeline we construct CineDub, the first Chinese television dubbing dataset with rich annotations. Across monologue, narration, dialogue, and complex multi-speaker scenes, Fun-CineForge consistently outperforms SOTA methods in audio quality, lip sync, timbre transfer, and instruction following — and is the first dubbing model to align long videos via the temporal modality. The model has been open-sourced through Alibaba Tongyi’s official speech repository (10K+ downloads), with releases on GitHub, Hugging Face, and ModelScope.

Publication
In Proceedings of IJCAI–ECAI 2026Oral
Jiaxuan Liu
Jiaxuan Liu
M.Eng. Student in Information & Communication Engineering · Speech AI Researcher

My research focuses on text-to-speech, expressive & emotional speech synthesis, multimodal movie dubbing, and speech foundation models.