返回全部动态
从经验中学习状态化预测知识
原标题:Learning Stateful Predictive Knowledge From Experience
AI 摘要
该论文提出状态知识学习(SKL)方法,使大语言模型代理从轨迹级反思转向维护显式的、基于状态的预测性知识。通过自蒸馏(SKL-SD)和强化学习(SKL-RL)两种算法,训练代理自主提取状态基础预测知识并用于决策。在WebShop、ScienceWorld和ChessPuzzles等交互环境中的实验表明,该方法显著优于当前基于反思的训练范式。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
Computer Science > Computation and Language Title:Learning Stateful Predictive Knowledge From Experience View PDF HTML (experimental) Abstract:As large language model (LLM) agents increasingly learn from experience, they primarily rely on trajectory-level reflection to extract insights. Viewed through the lens of predictive knowledge, we argue that this approach operates on episodic hindsight rather than predictive foresight, yielding brittle, path-dependent heuristics. To address th
发布时间:2026-08-03 12:00
抓取时间:2026-08-03 15:26
来源机构:arXiv