CSCR:反事实敏感度信用重分配改进长思维链推理
原标题:Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning
AI 摘要
该研究提出了一种名为CSCR(反事实敏感度信用重分配)的方法,用于改进长思维链推理中的强化学习。研究发现,在GRPO等无评论家方法中,特权信息引起的token似然偏移往往集中在可替换的表面形式token上,无法可靠指示正确的优化方向。CSCR通过减少对高敏感度token的信用并重新归一化token级优势,在数学推理基准上持续优于GRPO基线。
正文节选
Not All Tokens Deserve Equal Credit: Counterfactual Sensitivity Credit Reallocation for Long-CoT Reasoning Abstract Reinforcement learning with verifiable rewards (RLVR) is central to improving long-CoT reasoning in large language models. Critic-free methods such as GRPO convert response-level rewards into advantages and uniformly broadcast them across tokens, overlooking their unequal contributions to the final outcome. On-policy self-distillation (OPSD) instead provides dense distributional su