返回全部动态

CSCR:反事实敏感度信用重分配改进长思维链推理

原标题:Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning

Hugging Face Daily Papers一手来源研究质量 84

AI 摘要

该研究提出了一种名为CSCR(反事实敏感度信用重分配)的方法,用于改进长思维链推理中的强化学习。研究发现,在GRPO等无评论家方法中,特权信息引起的token似然偏移往往集中在可替换的表面形式token上,无法可靠指示正确的优化方向。CSCR通过减少对高敏感度token的信用并重新归一化token级优势,在数学推理基准上持续优于GRPO基线。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Abstract Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language models. Critic-free methods such as GRPO convert response-level rewards into advantages and uniformly broadcast them across tokens, overlooking their unequal contributions to the final outcome. On-policy self-distillation (OPSD) instead provides dense distributional su


发布时间:
抓取时间:2026-08-03 18:52
来源机构:Hugging Face
阅读原文huggingface.co