返回全部动态

参数空间探索通过变分学习改进 LLM 强化学习

原标题:Parameter Exploration for RLVR via Variational Learning

Hugging Face Daily Papers一手来源研究质量 84

AI 摘要

Hugging Face 每日论文介绍了一篇关于参数空间探索的研究,提出了一种名为 Perturbed Parameter Policy Optimization (3PO) 的方法族,通过从后验分布中采样不同策略来生成轨迹,以改进 LLM 的强化学习。实验表明,在 OLMo-3-1025-7B 和 Qwen2.5-Math-7B 上,3PO 在数学推理和代码生成任务中相比标准 GRPO 持续提升平均下游性能,且计算成本相近,同时减少了训练中的零优势组和错误轨迹。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Parameter Exploration for RLVR via Variational Learning Abstract Parameter-space exploration via perturbed policy sampling improves LLM reinforcement learning by diversifying rollouts and reducing training failures compared to action-space methods. Exploration has been a focus of reinforcement learning research for a long time. Recently, there has been growing evidence that it is also an important ingredient in LLM reinforcement learning recipes that can significantly impact downstream performan


发布时间:—
抓取时间:2026-08-13 18:50
来源机构:Hugging Face
阅读原文huggingface.co