Jiaxuan Liu is a Master’s student at the National Engineering Research Center of Speech and Language Information Processing (NERC-SLIP), University of Science and Technology of China. His research focuses on text-to-speech and speech foundation models, centering on highly expressive, human-like, emotionally rich and controllable speech synthesis, as well as multimodal movie dubbing and speech foundation models.
He is currently working with the Xiaomi MiMo team, contributing to data construction and pre-training for the next generation of speech foundation models on a corpus of hundreds of millions of hours of audio and video data.
He has published three first-author papers in top international venues. Fun-CineForge, which he developed, has been open-sourced through Alibaba Tongyi’s official speech repository and has received over 10,000 downloads.
M.Eng. in Information and Communication Engineering, 2024 – Present
University of Science and Technology of China (USTC), Hefei
B.Eng. in Computer Science and Technology (Rank 1/234), 2020 – 2024
Northwestern Polytechnical University (NWPU), Xi'an
Senior High School (Chemistry Olympiad Class), 2017 – 2020
Hengshui High School, Hebei
Open to collaboration in speech synthesis, multimodal generation, and speech foundation models — feel free to reach out.