返回全部动态

RLSVR:任务转换实现开放式 LLM 自改进的可验证奖励

原标题:From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement

Hugging Face Daily Papers一手来源研究质量 87

AI 摘要

Hugging Face 每日论文介绍了一项名为 RLSVR 的新训练范式,通过任务转换将开放式任务转化为可验证的代理环境,从而扩展 RLVR 的应用范围。该研究提出了 SpyRL 多智能体自博弈环境,在文本摘要、创意写作和数学推理等任务上表现出优于现有自改进方法的性能。相关模型和代码已在 GitHub 上发布。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

From RLVR to RLSVR: Task Transformation Induces Self-Verifiable Rewards for Open-Ended LLM Self-Improvement Abstract Reinforcement Learning with Verifiable Rewards (RLVR) has driven recent progress in reasoning-oriented large language models (LLMs) by enabling large-scale optimization. However, its applicability remains largely limited to domains such as mathematics and coding, where correctness can be deterministically verified. Open-ended tasks instead often rely on human preferences, reward m


发布时间:—
抓取时间:2026-08-03 16:11
来源机构:Hugging Face
阅读原文huggingface.co