Biography

Jiaxuan Liu is a Master’s student at the National Engineering Research Center of Speech and Language Information Processing (NERC-SLIP), University of Science and Technology of China. His research focuses on text-to-speech and speech foundation models, centering on highly expressive, human-like, emotionally rich and controllable speech synthesis, as well as multimodal movie dubbing and speech foundation models.

He is currently working with the Xiaomi MiMo team, contributing to data construction and pre-training for the next generation of speech foundation models on a corpus of hundreds of millions of hours of audio and video data.

He has published three first-author papers in top international venues. Fun-CineForge, which he developed, has been open-sourced through Alibaba Tongyi’s official speech repository and has received over 10,000 downloads.

Interests
  • Text-to-Speech (TTS) & Expressive Speech Synthesis
  • Emotional & Controllable Speech Generation
  • Multimodal Movie Dubbing
  • Speech Foundation Models / Speech LLM
  • Generative AI & Diffusion Models
Education
  • M.Eng. in Information and Communication Engineering, 2024 – Present

    University of Science and Technology of China (USTC), Hefei

  • B.Eng. in Computer Science and Technology (Rank 1/234), 2020 – 2024

    Northwestern Polytechnical University (NWPU), Xi'an

  • Senior High School (Chemistry Olympiad Class), 2017 – 2020

    Hengshui High School, Hebei

Education

 
 
 
 
 
University of Science and Technology of China · NERC-SLIP
M.Eng. in Information & Communication Engineering
September 2024 – Present Hefei, China
Research on text-to-speech, expressive & emotional synthesis, multimodal dubbing, and speech foundation models.
 
 
 
 
 
Northwestern Polytechnical University
B.Eng. in Computer Science & Technology
September 2020 – June 2024 Xi'an, China
Comprehensive score ranked 1 / 234 · Outstanding Graduate · recommended for admission to USTC.
 
 
 
 
 
Hengshui High School
Senior High School (Chemistry Olympiad Class)
July 2017 – July 2020 Hebei, China
First Prize, 33rd Chinese Chemistry Olympiad.

Research & Internship Experience

 
 
 
 
 
Xiaomi · MiMo Team
Speech Foundation Model Pre-training Intern
Xiaomi · MiMo Team
May 2026 – Present
Pre-training a speech foundation model on hundreds of millions of hours of data — multimodal understanding & unified audio synthesis.
 
 
 
 
 
Alibaba Cloud · Tongyi Lab
Algorithm Engineer Intern — Fun-CineForge
Alibaba Cloud · Tongyi Lab
July 2025 – May 2026
Led Fun-CineForge — a unified pipeline + multimodal model for zero-shot movie dubbing, plus the open-source CineDub dataset. Contributed to emotional-TTS training of CosyVoice3-0.5B. (IJCAI–ECAI 2026 Oral, 10K+ downloads.)
 
 
 
 
 
USTC · NERC-SLIP
NSFC Joint-Fund Project · UDDETTS
November 2024 – June 2025 Hefei, China
A unified discrete + dimensional (Arousal–Dominance–Valence) emotional-TTS framework with fine-grained linear 3-D controllability.
 
 
 
 
 
USTC · NERC-SLIP
NSFC Joint-Fund Project · DiffStyleTTS
June 2024 – October 2024 Hefei, China
Diffusion-based hierarchical prosody modeling for diverse and controllable TTS. (COLING 2025 Oral.)
 
 
 
 
 
iFLYTEK · Core R&D Platform
Research Intern · Undergraduate Thesis
December 2023 – April 2024 Hefei, China
Diffusion-based acoustic models for diverse and controllable prosodic TTS.
 
 
 
 
 
Huawei
Research Intern · Expressive Audiobook TTS
Huawei
June 2023 – July 2023
A multi-modal continuous emotional space + T5/contrastive text-emotion module to relieve audiobook listener fatigue.
 
 
 
 
 
ASLP@NPU Lab
Research Intern · Singing Voice Synthesis
ASLP@NPU Lab
June 2022 – July 2022 Xi'an, China
Contributed to the Opencpop dataset; optimised ByteSing’s Duration / LF0 / U-V Predictor with BLSTM. Outstanding Intern of ASLP Lab.

Honors & Awards

🎓 Scholarships

  • Zenghua Special Scholarship, USTC · 2025
  • USTC 0006 Scholarship · 2025
  • Master’s First-Class Academic Scholarship, USTC · 2024, 2025
  • Xiaomi Excellent Scholarship · 2023
  • National Inspirational Scholarship · 2022, 2023
  • Yajun Wu Special Scholarship, NWPU · 2022
  • First-Class Academic Scholarship, NWPU · 2021, 2022, 2023
  • National Scholarship · 2021

🏆 Competitions (Selected)

  • Mathematical Contest in Modeling (MCM) — Outstanding Winner · 2023
  • CUMCM — Shaanxi First Prize · 2022
  • CUPT — Shaanxi Individual & Team First Prize · 2021
  • Baidu Wenxin LLM Creative Application — A- Award · 2022
  • “Challenge Cup” Shaanxi Entrepreneurship Plan — Bronze · 2022
  • NWPU Programming / Mathematics / Physics Experiment — First Prize

✨ Other Honors

  • Outstanding Communist Party Member, USTC · 2025, 2026
  • Granted Chinese Invention Patent — Lightweight End-to-End Object Segmentation for Road-Damage Detection · 2025
  • Outstanding Graduate of NWPU · 2024
  • Outstanding Student of NWPU · 2021, 2022, 2023
  • Outstanding Communist Youth League Member (May 4th Honors) · 2021, 2022
  • Outstanding Intern, ASLP@NPU Lab · 2022
  • 14th National Games of the PRC — Flag Phalanx, commended by the State Council & CYLC Central Committee · 2021

Service & Activities

  • Secretary, 2nd Postgraduate Party Branch, Class of 2024, School of Information Science and Technology, USTC · 2024 – Present
  • Class Study Committee Member, NWPU · 2020 – 2024

Contact

Open to collaboration in speech synthesis, multimodal generation, and speech foundation models — feel free to reach out.