返回全部动态

从失语症命名错误谱逆向恢复大语言模型损伤参数

原标题:Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models

arXiv cs.CL一手来源研究质量 83

AI 摘要

该研究提出从失语症图片命名错误谱中逆向恢复大型语言模型(LLaVA-Vicuna 13B)的损伤参数(层索引、修改百分比、噪声sigma),并训练多任务神经网络实现映射。结果显示修改百分比和噪声sigma可恢复,层索引仅能近似恢复,但反事实验证中恢复参数能高保真复现目标行为(81.4%),表明Transformer层存在功能冗余。在278名中风幸存者的错误谱上,恢复参数具有综合征区分性,验证了跨分布泛化能力。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Computer Science > Computation and Language Title:Recovering Lesion Parameters from Aphasic Picture Naming Error Profiles in Large Language Models View PDF Abstract:Interpretability methods for large language models (LLMs) describe internal state but do not directly test whether that state is causally sufficient to produce the observed behavior. In earlier work, we lesioned LLMs to produce error profiles in picture naming, a central task for assessing aphasia, and found that specific


发布时间:2026-08-10 12:00
抓取时间:2026-08-10 12:01
来源机构:arXiv
阅读原文arxiv.org