返回全部动态
ZipFormer-30M 越南语语音识别模型发布
原标题:hynt/Zipformer-30M-RNNT-6000h
AI 摘要
Hugging Face 上发布了越南语语音识别模型 ZipFormer-30M-RNNT-6000h,基于 ZipFormer 架构,仅 30M 参数,在 CPU 上转录 12 秒音频仅需 0.3 秒。该模型在 VLSP 2025 竞赛中获第一名,并在多个基准上优于更大模型。模型采用 RNN-T 损失,训练数据约 6000 小时,支持快速 CPU 推理,适合实时语音识别等场景。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
--- license: cc-by-nc-nd-4.0 --- # Vietnamese Speech-to-Text (ASR) — ZipFormer-30M-RNNT-6000h ## 🔍 Overview The **Vietnamese Speech-to-Text (ASR)** model is built on the **ZipFormer architecture** — an improved variant of the Conformer — featuring only **30 million parameters** yet achieving **exceptional performance** in both speed and accuracy. On CPU, the model can transcribe a **12-second audio clip in just 0.3 seconds**, significantly faster than most traditional ASR systems without requ
发布时间:2026-09-09 18:13
抓取时间:2026-09-09 18:14
来源机构:Hugging Face