返回全部动态
ERR+:基于序列熵消解的 LLM 高效推理框架
原标题:ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning
AI 摘要
ERR+ 提出一种两阶段 RLVR 框架,基于 token 级熵降奖励(ERR)和鲁棒相对效率奖励(RRER)优化 LLM 推理。实验表明正确推理轨迹熵降更频繁,ERR+ 在五个数据集上提升准确率并缩短响应长度。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
ERR+: Sequential Entropy Resolution for Efficient and Decisive LLM Reasoning Abstract Large reasoning models achieve strong performance on complex tasks by generating extended chain-of-thought (CoT) traces via reinforcement learning with verifiable rewards (RLVR). While current RLVR methods have achieved strong results with correctness-based reward signals, they provide limited guidance on the quality of the reasoning process itself, leaving the internal reasoning structure largely unoptimized.
发布时间:2026-09-01 12:00
抓取时间:2026-09-01 12:13
来源机构:arXiv