Xingchen Song 宋星辰
I work on speech recognition, speech synthesis, and audio language models, with a focus on efficient training and inference. I develop open-source tools for multimodal model training, speech tokenization, and speech generation.
Research Interests
- Streaming and end-to-end speech recognition
- Speech synthesis and audio language models
- Efficient multimodal training and inference
Selected Publications
For the complete list, please visit my Google Scholar profile.
2026
- Semantic Refinement of Universal Audio Representations through Audio-Description Alignment
- Qwen-Audio-3.0-Gen-Preview Technical Report
- Harness TTS: Towards Context-Aware Expressive Speech Synthesis with Harness Layer
- F3-Tokenizer: Taming Audio Autoencoder Latents for Understanding and Generation
- Any2Speech: Borderless Long Speech Synthesis
- ZipVoice-Dialog: Non-Autoregressive Spoken Dialogue Generation with Flow Matching
2025
- MiMo-Audio: Audio Language Models are Few-Shot Learners
- Adapting Whisper for Streaming Speech Recognition via Two-Pass Decoding
2024
- TouchASP: Elastic Automatic Speech Perception that Everyone Can Touch
- TouchTTS: An Embarrassingly Simple TTS Framework that Everyone Can Touch
- U2++ MoE: Scaling 4.7x parameters with minimal impact on RTF
2023
- ZeroPrompt: Streaming Acoustic Encoders are Zero-Shot Masked LMs
- TrimTail: Low-Latency Streaming ASR with Simple But Effective Spectrogram-Level Length Penalty
2022
2021
2020
Open-Source Projects
-
TouchNet
Personal project · In development
A native PyTorch library for large-scale text and audio language model training, with multimodal data pipelines, sequence packing, distributed checkpointing, and N-dimensional parallelism.
-
FlashCosyVoice
Personal project
A lightweight inference engine for CosyVoice, built for distributed offline speech generation. It supports prefix caching, torch compilation, and CUDA graphs in a compact Python implementation.
-
S3Tokenizer
Personal project
A PyTorch implementation of the supervised semantic speech tokenizer used in CosyVoice, supporting batch and distributed inference, online speech token extraction, and long audio processing.
-
WeNet
Contributor
A production-oriented toolkit for streaming and non-streaming end-to-end speech recognition. My related publications include WeNet 2.0, TrimTail, ZeroPrompt, and U2++ MoE.
-
QwenAudio Toolkits
Contributor
A desktop workspace for audio AI models, covering speech recognition, audio enhancement, text normalization, and speech synthesis, with local model runtimes and optional cloud models.
Education
-
M.S. in Computer Science and Technology
Tsinghua University · 2019–2022
Advisor: Zhiyong Wu
-
B.E. in Computer Science and Technology
Dalian University of Technology · 2015–2019
Experience
-
Research Intern @ Microsoft Research Asia (MSRA)
Speech Group · Dec. 2020–Mar. 2021
Mispronunciation detection and diagnosis.
Mentors: Frank K. Soong
-
Research Intern @ Tencent AI Lab
Speech Group · Jun. 2019–Dec. 2020
Acoustic modeling and end-to-end speech recognition.
Mentors: Guangsen Wang, Yiheng Huang