返回全部动态

推荐基础模型的多阶段后训练渐进对齐框架

原标题:Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training

arXiv cs.IR一手来源研究质量 82

AI 摘要

该论文提出了一种用于推荐基础模型的三阶段渐进式后训练框架,将下游适配与业务指标对齐分离。适配阶段包括线性探测和全量微调,随后通过强化学习微调(RFT)使用学习到的奖励模型对齐业务目标。离线实验和在线A/B测试表明,该框架优于单阶段替代方案,并提升了生产推荐质量。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Computer Science > Information Retrieval Title:Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training View PDF HTML (experimental) Abstract:Foundation model(FM) for recommendation has shown strong ability to model long-horizon sequential user behavior. In practice, a single pretrained foundation model is often adapted to diverse downstream serving surfaces through Supervised Fine-Tuning(SFT). However, optimizing task-specific objectives such as clicks


发布时间:2026-08-10 12:00
抓取时间:2026-08-10 13:04
来源机构:arXiv
阅读原文arxiv.org