返回全部动态

分离语音语言模型中的决策规则错位与读出覆盖限制

原标题:Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models

arXiv cs.CL一手来源研究质量 81

AI 摘要

该研究提出了一种生成对齐的诊断阶梯方法,用于分离语音语言模型在副语言任务中决策规则与读出覆盖范围的性能差距。实验发现,状态解码平均比生成准确率高27.8个百分点,且标签无关的logit修正能改善生成准确性。结果表明,情感信息在原生读出之外可泛化到未见说话者,但替换读出外部方向对生成答案影响甚微。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Computer Science > Computation and Language Title:Separating Decision-Rule Misalignment from Readout-Coverage Limitations in Speech Language Models View PDF HTML (experimental) Abstract:Speech language models are increasingly evaluated on paralinguistic tasks by the accuracy of prompted answers, but answer accuracy combines failures at different stages of the audio-to-answer computation. We introduce a generation-aligned diagnostic ladder that compares the emitted answer, the option


发布时间:2026-08-10 12:00
抓取时间:2026-08-10 12:01
来源机构:arXiv
阅读原文arxiv.org