项目Projects

53404
参与项目 GitHub StarsGitHub stars across my projects
6 个参与项目 · 含一作 2 项与非一作 4 项6 projects I contributed to — 2 as first author, 4 as a contributing author
799866
Hugging Face 下载量Hugging Face downloads
9 个模型与数据集 · 近 30 天9 models and datasets — trailing 30 days

数据来自 GitHub 与 Hugging Face API,页面打开时实时获取;获取失败时回落到 2026 年 8 月 15 日 的快照。Fetched live from the GitHub and Hugging Face APIs; falls back to a snapshot taken 15 August 2026 if that fails.

2026-09
microsoft/VibeVoice 一作First author
开源前沿语音 AI —— VibeVoice-ASR-Streaming 是我的一作工作,最早一批基于大模型的端到端流式 speaker-attributed ASR 之一,边说边输出「谁说了什么」。同一仓库还包含我参与的 VibeVoice-ASR,支持单次处理长达 60 分钟的音频。Open-Source Frontier Voice AI — VibeVoice-ASR-Streaming, my first-author work, is one of the first LLM-based end-to-end approaches to streaming speaker-attributed ASR, emitting "who said what" as speech arrives. The same repository hosts VibeVoice-ASR, which I contributed to, for single-pass processing of up to 60 minutes of audio.
GitHub stars GitHub forks arXiv Hugging Face model microsoft/VibeVoice-ASR-Streaming-7B Hugging Face model microsoft/VibeVoice-ASR-Streaming-1.5B Hugging Face model microsoft/VibeVoice-ASR Hugging Face model microsoft/VibeVoice-ASR-HF
2026-07
microsoft/VibeASR.cpp 参与Contributor
VibeVoice-ASR-BitNet —— VibeVoice-ASR 的压缩版本,用 INT8 声学 token 化与三值语言模型权重,在边缘 CPU 上实现实时多语种语音识别。VibeVoice-ASR-BitNet — a compressed VibeVoice-ASR variant for real-time multilingual speech recognition on edge CPUs, using INT8 acoustic tokenization and ternary language-model weights.
GitHub stars GitHub forks arXiv Hugging Face model microsoft/VibeVoice-ASR-BitNet
2026-06
SpeechColab/GigaSpeechBench 一作First author
GigaSpeechBench —— 我的一作工作,一个 680 小时的真实场景多语种 ASR 与 AST 评测基准,覆盖低资源语言、方言、口音、垂直领域与不同年龄人群。GigaSpeechBench — my first-author work introducing a 680-hour, real-world multilingual ASR & AST benchmark spanning low-resource languages, dialects, accents, domains, and age groups.
GitHub stars GitHub forks arXiv Hugging Face dataset speechcolab/GigaSpeechBench
2026-05
Gilgamesh-J/X-ASR 参与Contributor
X-ASR —— 面向流式的语音识别模型系列。首个版本是 160M 参数的中英 Zipformer transducer,在约一百万小时语音上训练,统一了离线与真流式识别,并支持 sherpa-onnx 部署。X-ASR — a series of streaming-focused automatic speech recognition models. Its first release is a 160M-parameter Chinese-English Zipformer transducer trained on approximately one million hours of speech, unifying offline and true streaming recognition with sherpa-onnx deployment.
GitHub stars GitHub forks Hugging Face model GilgameshWind/X-ASR-zh-en Technical report coming soon
2026-05
sjtuayj/OpenSTBench 参与Contributor
OpenSTBench —— 面向语音翻译的统一多维度评测框架,覆盖离线与流式设定下的 S2TT 与 S2ST 系统,联合衡量翻译质量、语音质量、说话人保持、情感与副语言信息还原、时序一致性与延迟。OpenSTBench — a unified multidimensional evaluation framework for speech translation, covering S2TT and S2ST systems in both offline and streaming settings. It jointly measures translation quality, speech quality, speaker preservation, emotion and paralinguistic fidelity, temporal consistency, and latency.
GitHub stars GitHub forks arXiv Hugging Face dataset ayj111/openstbench-paired-set
2026-01
SWivid/Habibi-TTS 参与Contributor
Habibi —— 首个开源的统一多方言阿拉伯语语音合成框架,覆盖 12 种以上地区方言,效果可与 ElevenLabs Eleven v3 (alpha) 相比。Habibi — the first open-source unified-dialectal Arabic TTS framework, covering 12+ regional dialects. Competitive with ElevenLabs Eleven v3 (alpha).
GitHub stars GitHub forks arXiv Hugging Face model SWivid/Habibi-TTS