NCP-Bench:评估大语言模型在交互叙事中的长程一致性
原标题:Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives
AI 摘要
Hugging Face 每日论文介绍了一项新研究,提出了叙事承诺保持(NCP)概念和 NCP-Bench 基准,用于评估大语言模型在交互式叙事中的长程逻辑一致性。基准包含 100 个基于电影情节的叙事环境,实验发现即使最强模型 GPT-5.2 在 20 轮后也只有 42% 的存活率,事实冲突率在 40% 到 68% 之间,表明当前模型在长程交互中难以保持一致性。
正文节选
Can LLM Agents Stick to the Script? A Benchmark for Long-Horizon Consistency in Interactive Narratives Abstract The study introduces a benchmark and formalizes narrative commitment preservation to evaluate long-horizon logical consistency in interactive storytelling with large language models. The rapid advancement of Large Language Models (LLMs) is revolutionizing AI for Games by enabling open-ended and fluid interactive storytelling. However, existing research has largely overlooked the critic