Google 提出 Retrieve-for-Train:用离线 RL 加速复杂 AI 搜索
原标题:Bypassing inference bottlenecks: Accelerating complex AI search with Retrieve-for-Train
AI 摘要
Google Research 在 ICML 2026 论文中提出 Retrieve-for-Train 框架,通过离线强化学习将奖励对齐的查询扇出行为编译为监督信号,再蒸馏进轻量扩散检索器,从而在推理时单次生成连贯、专家级的搜索结果集合。该方法在 Gemma3-4B 和 Qwen3-4B 上使用 Soft-GRPO 训练,并在时尚文本到图像与音乐文本到音乐两个集合级检索任务上超越传统方法,避免推理时大量思考预算开销。
正文节选
September 15, 2026 Pengcheng Jiang, Student Researcher, and Judith Yue Li, Senior Research Engineer, Google Research Instead of relying on expensive inference-time reasoning, the Retrieve-for-Train framework uses reinforcement learning once to train a lightweight diffusion model. This bypasses the heavy autoregressive "thinking budget" to instantly generate a cohesive, expert-level slate of AI search results. Modern search or recommendation applications are increasingly expected to return a cohe