LongNovel:长上下文小说摘要幻觉检测的多尺度基准
原标题:LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization
AI 摘要
华东师范大学、爱奇艺和腾讯的研究者提出了 LongNovel,一个用于长上下文小说摘要幻觉检测的多尺度中英双语基准。该基准包含 29 部中文小说和 BookSum 数据集的章节级数据,覆盖 16k 到 100k token 的四种长度场景,并设计了 8 种幻觉类型。通过多模型仲裁和实体引用幻觉生成相结合的方法,确保数据真实性和类别平衡,实验表明 LongNovel 具有挑战性,已开源供研究使用。
正文节选
LongNovel: A Multi-Scale Benchmark for Hallucination Detection in Long-Context Novel Summarization Abstract Although context windows have expanded significantly in recent years, hallucinations in long-context summarization remain a challenge. Long novels are better suited than news or papers for researching these hallucinations, due to their intrinsic information and detailed descriptions of events and dialogues. However, current research lacks a multi-scale benchmark for hallucination detection