返回全部动态
推荐基础模型的多阶段后训练渐进对齐框架
原标题:Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training
AI 摘要
该论文提出了一种用于推荐基础模型的三阶段渐进式后训练框架,将下游适配与业务指标对齐分离。适配阶段包括线性探测和全量微调,随后通过强化学习微调(RFT)使用学习到的奖励模型对齐业务目标。离线实验和在线A/B测试表明,该框架优于单阶段替代方案,并提升了生产推荐质量。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
Computer Science > Information Retrieval Title:Progressive Alignment of Recommender Foundation Model through Multi-Phase Post-Training View PDF HTML (experimental) Abstract:Foundation model(FM) for recommendation has shown strong ability to model long-horizon sequential user behavior. In practice, a single pretrained foundation model is often adapted to diverse downstream serving surfaces through Supervised Fine-Tuning(SFT). However, optimizing task-specific objectives such as clicks
发布时间:2026-08-10 12:00
抓取时间:2026-08-10 13:04
来源机构:arXiv