返回全部动态
低频输入引发模型故障:LALMs中的安全风险研究
原标题:From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs
AI 摘要
研究人员提出了一种名为ILL的黑盒红队方法,利用听不见的低频波形暴露音频语言模型的漏洞,并设计了DRG防御机制来检测分布偏移并恢复准确性。实验表明,ILL可使模型准确率最多降低67个百分点,而DRG能将受攻击后的平均准确率从28.5%提升至46.1%。该研究揭示了LALMs中此前被忽视的安全风险。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
From Inaudible Inputs to Model Failures: Low-Frequency Safety Risks in LALMs Abstract Researchers propose a black-box red-teaming method using inaudible low-frequency waveforms to expose vulnerabilities in audio-language models, alongside a defense that detects distribution shifts and requests a second recording to recover accuracy. Large audio-language models (LALMs) have demonstrated strong capabilities in understanding diverse audio inputs. This diversity includes low-frequency signals that a
发布时间:—
抓取时间:2026-08-14 23:19
来源机构:Hugging Face