OpenSTBench: Beyond Semantic Evaluation for Speech Translation

发表于Published in arXiv preprint arXiv:2605.30792, 2026

Speech translation systems increasingly span speech-to-text translation (S2TT), speech-to-speech translation (S2ST), offline translation, and streaming generation, producing outputs that differ in modality, speech realization, and timing behavior. Existing evaluation practices assess translation quality, speech quality, and temporal quality under separate protocols, which makes heterogeneous systems difficult to compare. OpenSTBench organizes these outputs into a shared evaluation format and jointly evaluates translation quality, speech quality, speaker preservation, emotion and paralinguistic fidelity, temporal consistency, and latency. Experiments on representative systems show that strong translation quality does not imply strong speech or temporal quality, and the framework provides a reproducible protocol for analyzing those cross-dimensional differences.

GitHub stars

Paper Code

推荐引用:Recommended citation: Y. An, Y. Zhao, Y. Zhang, Q. Zheng, Y. Tu, K. Deng, K. Yu, and X. Chen. (2026). "OpenSTBench: Beyond Semantic Evaluation for Speech Translation." arXiv preprint arXiv:2605.30792.
下载论文Download Paper