论文Publications

也可以在 我的 Google Scholar 主页查看这些工作。You can also find my articles on my Google Scholar profile.

低资源语音技术Low-Resource Speech Technologies

为常规评测长期忽视的语言、方言与人群,构建基准数据与识别、合成等语音技术。Benchmarks, recognition and synthesis for languages, dialects, and communities that conventional evaluation leaves out.

2 works

GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

发表于Published in arXiv preprint arXiv:2606.28884, 2026

First-author work. GigaSpeechBench is a 680-hour, human-annotated, real-world ASR & AST benchmark covering low-resource languages, Chinese dialects, English accents, vertical-domain terminology, and speech from older adults and children.
GitHub stars

推荐引用:Recommended citation: Y. Tu, Y. Yang, T. Wang, Y. Zhu, G. Lin, M. Shao, H. Wang, J. Liu, Y. Fu, et al. (2026). "GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark." arXiv preprint arXiv:2606.28884.
下载论文Download Paper

Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis

发表于Published in arXiv preprint arXiv:2601.13802, 2026

Habibi is a unified-dialectal Arabic TTS framework covering 12+ regional dialects. Our unified model matches or surpasses per-dialect specialized models and is highly competitive with ElevenLabs Eleven v3 (alpha).
GitHub stars

推荐引用:Recommended citation: Y. Chen, J. Liu, Y. Tu, Z. Niu, Y. Liang, C. Qiang, C. Zhang, K. Yu, X. Chen. (2026). "Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis." arXiv preprint arXiv:2601.13802.
下载论文Download Paper

多说话人语音识别Multi-Speaker Speech Recognition

把转写、说话人归属与实际部署统一起来的长音频高效语音识别系统。Long-form and efficient ASR systems that unify transcription, speaker attribution, and practical deployment.

2 works

VibeVoice-ASR-BitNet Technical Report

发表于Published in arXiv preprint arXiv:2607.21075, 2026

VibeVoice-ASR-BitNet compresses VibeVoice-ASR for real-time multilingual recognition on edge CPUs through heterogeneous quantization, custom SIMD kernels, and fused operators. It is 1.6-2.3x faster than Whisper.cpp at comparable model sizes.
GitHub stars

推荐引用:Recommended citation: S. Xu, T. Song, S. Huang, Z. Peng, Y. Xia, Y. Tu, X. Huang, X. Wu, W. Wang, Y. Chang, J. Yu, L. Dong, and F. Wei. (2026). "VibeVoice-ASR-BitNet Technical Report." arXiv preprint arXiv:2607.21075.
下载论文Download Paper

VIBEVOICE-ASR Technical Report

发表于Published in arXiv preprint arXiv:2601.18184, 2026

VibeVoice-ASR is a general-purpose speech understanding framework that supports single-pass processing for up to 60 minutes of audio, unifying ASR, Speaker Diarization, and Timestamping into a single end-to-end generation task. It supports over 50 languages and natively handles code-switching.
GitHub stars

推荐引用:Recommended citation: Z. Peng, J. Yu, Y. Chang, Z. Wang, L. Dong, Y. Hao, Y. Tu, C. Yang, W. Wang, et al. (2026). "VIBEVOICE-ASR Technical Report." arXiv preprint arXiv:2601.18184.
下载论文Download Paper

语音合成、翻译和其他语音技术Speech Synthesis, Translation and Other Speech Technologies

语音翻译与合成系统,以及它们所需的评测方法,兼及相邻的语音技术方向。Speech translation and synthesis systems, and the evaluation protocols they need, alongside neighbouring speech technologies.

1 works

OpenSTBench: Beyond Semantic Evaluation for Speech Translation

发表于Published in arXiv preprint arXiv:2605.30792, 2026

OpenSTBench is a unified multidimensional evaluation framework for speech translation, covering S2TT and S2ST systems in both offline and streaming settings. It jointly measures translation quality, speech quality, speaker preservation, emotion and paralinguistic fidelity, temporal consistency, and latency.
GitHub stars

推荐引用:Recommended citation: Y. An, Y. Zhao, Y. Zhang, Q. Zheng, Y. Tu, K. Deng, K. Yu, and X. Chen. (2026). "OpenSTBench: Beyond Semantic Evaluation for Speech Translation." arXiv preprint arXiv:2605.30792.
下载论文Download Paper