GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark

Published in arXiv preprint arXiv:2606.28884, 2026

GigaSpeechBench is a comprehensive multilingual and multidimensional in-the-wild ASR and AST benchmark comprising 680 hours of human-annotated speech. Its five modules cover low-resource Middle Eastern and Southeast Asian languages plus Japanese and Korean, Chinese dialects, English accents, dense terminology across 12 vertical domains, and speech from older adults and children. Human-annotated Chinese and English translations for 11 languages also support speech translation evaluation.

GitHub stars

Paper Code Dataset

Recommended citation: Y. Tu, Y. Yang, T. Wang, Y. Zhu, G. Lin, M. Shao, H. Wang, J. Liu, Y. Fu, et al. (2026). "GigaSpeechBench: A Real-World Multilingual Speech-to-Text Benchmark." arXiv preprint arXiv:2606.28884.
Download Paper