RA-GRPO:基于反思的偏好优化用于视觉生成
原标题:Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation
AI 摘要
本文提出了一种名为 RA-GRPO 的强化学习偏好对齐框架,用于扩散生成模型。该方法通过引入 Diffusion Reflection 机制,利用逆向扩散过程修正中间采样轨迹,并借助 Counterfactual Path Synthesis 将修正后的轨迹蒸馏到策略中,以提升文本到图像和文本到视频生成的语义保真度和视觉真实感。实验表明,RA-GRPO 在缓解奖励黑客和改善泛化方面优于现有方法,且与标准流程无缝集成。
正文节选
Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation Abstract. Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advanced multimedia content synthesis, especially in text-to-image and text-to-video tasks. To further align such generative models with human preferences, reinforcement learning (RL) has recently shown strong potential as a post-training strategy. Nevertheless, existing policy gradient-based m