返回全部动态

通过注意力图拓扑特征检测大语言模型幻觉

原标题:Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing

arXiv cs.AI一手来源研究质量 86

AI 摘要

该研究提出通过分析注意力图的拓扑结构来区分大语言模型的幻觉与非幻觉回答,利用Forman-Ricci曲率识别注意力图中的信息瓶颈。作者提出一组新的拓扑特征,使简单线性探针在两个幻觉检测基准上持续优于基于注意力和多响应基线,并在三种LLM家族上验证。分析发现幻觉回答与因果生成过程中token间上下文共享受损密切相关,尤其表现为过度依赖自注意力、上下文检索分散或信息过度压缩,且集中在最后一层Transformer。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Detecting Hallucination in LLMs: Tracing the Topological Signatures of Impaired Context Sharing Abstract In this work, we examine the topology of information flow patterns within attention graphs to effectively distinguish hallucinated from non-hallucinated responses. We analyze the Forman–Ricci curvature to identify structural patterns indicating information bottlenecks in attention graphs. We then introduce a method that captures both semi-local and global information-flow characteristics of a


发布时间:2026-09-22 12:00
抓取时间:2026-09-21 12:13
来源机构:arXiv
阅读原文arxiv.org