Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis

发表于Published in arXiv preprint arXiv:2601.13802, 2026

Figure from Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis

Arabic spans over 30 spoken varieties, yet no open-source text-to-speech system unifies them. We present Habibi, a unified-dialectal Arabic TTS framework that addresses key barriers including substantial cross-dialect lexical and phonological divergence, scarce synthesis-grade data, and the absence of a standardized multi-dialect evaluation benchmark. Both automatic metrics and human evaluations confirm that Habibi is highly competitive with ElevenLabs’ Eleven v3 (alpha) in intelligibility, speaker similarity, and naturalness.

GitHub stars

Paper Code Demo

推荐引用:Recommended citation: Y. Chen, J. Liu, Y. Tu, Z. Niu, Y. Liang, C. Qiang, C. Zhang, K. Yu, X. Chen. (2026). "Habibi: Laying the Open-Source Foundation of Unified-Dialectal Arabic Speech Synthesis." arXiv preprint arXiv:2601.13802.
下载论文Download Paper