I work in the field of Audio Singal Processing, Speech Representation, and Multimodal Large Language Model supervised by Assoc. Prof. Xie Chen, I will try my best in the next five exciting years! 💪. I am also interested in building scalable systems. Currently, I focus on the following research topics:
Jeongsoo Choi*, Zhikang Niu*, Ji-Hoon Kim, Chunhui Wang, Joon Son Chung, Xie Chen. Accelerating Diffusion-based Text-to-Speech Model Training with Dual Modality Alignment. InterSpeech 2025 Oral,
[Link]
[PDF]
[Code]
[BibTeX]
Yushen Chen, Zhikang Niu, Ziyang Ma, Keqi Deng, Chunhui Wang, Jian Zhao, Kai Yu, Xie Chen. F5-TTS: A Fairytaler that Fakes Fluent and Faithful Speech with Flow Matching. ACL 2025 Main,
[Link]
[PDF]
[Code]
[BibTeX]
[Talk] F5-TTS has collected 13,000+ stars on GitHub.
Wenhao Guan, Zhikang Niu, Ziyue Jiang, Kaidi Wang, Peijie Chen, Qingyang Hong, Lin Li, Xie Chen. UniVoice: Unifying Autoregressive ASR and Flow-Matching based TTS with Large Language Models. InterSpeech 2026,
[Link]
[PDF]
[Code]
[BibTeX]
MMAR Team. MMAR: A Challenging Benchmark for Deep Reasoning in Speech, Audio, Music, and Their Mix. NeurIPS Dataset & Benchmark 2025,
[Link]
[PDF]
[Code]
thorough-pytorch: A Chinese PyTorch tutorial and it has already collected 2,300 more stars and 333 forks on GitHub.
CSBasicKnowledge: This repo will record some knowledge about computer science, artificial intelligence and EE. It has already collected 560 more stars on GitHub.
More open-source contents can be found on my GitHub.
2021.11-Now, Datawhale member (an open-source AI organization), helped data science fans get involved in the AI community.
2021.11-Now, Xmart forum maintainer (an open-source student forum from SJTU X-LANCE Lab), for helping students get involved in the speech AI community.