返回全部动态

SSR-GRPO:融合监督与语义标识符的电商密集检索强化学习

原标题:SSR-GRPO: Integrating Supervision and Semantic IDs into Reinforcement Learning for Dense Retrieval in E-commerce

arXiv cs.IR一手来源研究质量 84

AI 摘要

本文提出 SSR-GRPO,一种用于电商密集检索的强化学习框架,通过引入语义标识符(SIDs)和双视角相关性评估,解决 R-GRPO 中候选池噪声和奖励偏差问题。该方法利用 SIDs 挖掘难负样本,并集成掩码函数和 Retrieval-DPO 任务,提升模型细粒度语义区分能力。离线与在线实验验证了其有效性,并已部署于大规模电商平台。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

SSR-GRPO: Integrating Supervision and Semantic IDs into Reinforcement Learning for Dense Retrieval in E-commerceDOI: 10.1145/3799682.3840101Conference: Proceedings of the 35th ACM International Conference on Information and Knowledge Management; November 07–11, 2026; Rome, ItalyProceedings of the 35th ACM International Conference on Information and Knowledge Management (CIKM ’26), November 07–11, 2026, Rome, ItalyISBN: 979-8-4007-2539-5/2026/11CCS: Information systems Retrieval models and rankin


发布时间:2026-08-21 12:00
抓取时间:2026-08-21 12:32
来源机构:arXiv
阅读原文arxiv.org