RLSVR:任务转换实现开放式 LLM 自改进的可验证奖励
原标题:From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement
AI 摘要
Hugging Face 每日论文介绍了一项名为 RLSVR 的新训练范式,通过任务转换将开放式任务转化为可验证的代理环境,从而扩展 RLVR 的应用范围。该研究提出了 SpyRL 多智能体自博弈环境,在文本摘要、创意写作和数学推理等任务上表现出优于现有自改进方法的性能。相关模型和代码已在 GitHub 上发布。
正文节选
From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Abstract Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verified. Open-ended tasks instead often rely on human preferences, reward m