返回全部动态

REACH:通过LLM增强的音频-文本对齐实现零样本呼吸音分类

原标题:Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment

arXiv cs.CL一手来源研究质量 83

AI 摘要

REACH框架通过LLM增强的音频-文本对齐,将自监督呼吸编码器转化为零样本能力的基础模型,在6个数据集的9项任务中平均零样本AUC达61.3%,超越CLAP和Qwen2-Audio,同时线性探测AUC达71.6%,仅使用全规模基线43%的数据。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

İlerisoy Pham Funk Pechenizkiy Saeed Zero-Shot Respiratory Sound Classification through LLM-Augmented Audio-Text Alignment Abstract Self-supervised respiratory encoders lack semantic grounding in clinical domain needed for zero-shot inference, limiting their utility without task-specific labeled data. We propose a framework that aligns these encoders with medical terminology in a shared latent space turning them into a zero-shot-capable foundation model. To address paired data scarcity, we use a


发布时间:2026-09-02 12:00
抓取时间:2026-09-02 12:10
来源机构:arXiv
阅读原文arxiv.org