返回全部动态

RA-GRPO:基于反思的偏好优化用于视觉生成

原标题:Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation

arXiv cs.CV一手来源研究质量 87

AI 摘要

本文提出了一种名为 RA-GRPO 的强化学习偏好对齐框架,用于扩散生成模型。该方法通过引入 Diffusion Reflection 机制,利用逆向扩散过程修正中间采样轨迹,并借助 Counterfactual Path Synthesis 将修正后的轨迹蒸馏到策略中,以提升文本到图像和文本到视频生成的语义保真度和视觉真实感。实验表明,RA-GRPO 在缓解奖励黑客和改善泛化方面优于现有方法,且与标准流程无缝集成。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation Abstract. Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advanced multimedia content synthesis, especially in text-to-image and text-to-video tasks. To further align such generative models with human preferences, reinforcement learning (RL) has recently shown strong potential as a post-training strategy. Nevertheless, existing policy gradient-based m


发布时间:2026-09-07 12:00
抓取时间:2026-09-07 12:03
来源机构:arXiv
阅读原文arxiv.org