SSR-GRPO:融合监督与语义标识符的电商密集检索强化学习
原标题:SSR-GRPO: Integrating Supervision and Semantic IDs into Reinforcement Learning for Dense Retrieval in E-commerce
AI 摘要
本文提出 SSR-GRPO,一种用于电商密集检索的强化学习框架,通过引入语义标识符(SIDs)和双视角相关性评估,解决 R-GRPO 中候选池噪声和奖励偏差问题。该方法利用 SIDs 挖掘难负样本,并集成掩码函数和 Retrieval-DPO 任务,提升模型细粒度语义区分能力。离线与在线实验验证了其有效性,并已部署于大规模电商平台。
正文节选
SSR-GRPO: Integrating Supervision and Semantic IDs into Reinforcement Learning for Dense Retrieval in E-commerceDOI: 10.1145/3799682.3840101Conference: Proceedings of the 35th ACM International Conference on Information and Knowledge Management; November 07–11, 2026; Rome, ItalyProceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM ’26), November 07–11, 2026, Rome, ItalyISBN: 979-8-4007-2539-5/2026/11CCS: Information systems Retrieval models and rankin