Search-G1:基于表征内在奖励的接地搜索代理
原标题:Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards
AI 摘要
arXiv 论文提出 Search-G1,一种基于表征的内在奖励框架,用于训练搜索增强语言代理。该框架通过两个干预校准的读数(提示状态读数和答案承诺读数)衡量代理答案的操作性接地,以区分必要检索与冗余搜索,并在强化学习过程中周期性地重新拟合读数,使奖励与策略共同演化。实验表明,Search-G1 在多个基于搜索的问答基准上改善了接地与搜索成本的权衡,在保持竞争性任务准确率的同时缩短了响应侧轨迹。
正文节选
Computer Science > Computation and Language Title:Search-G1: Grounded Search Agents via Representation-Based Intrinsic Rewards View PDF HTML (experimental) Abstract:Search-augmented language agents should retrieve external information only when necessary and ground their answers in retrieved evidence. Existing external rewards provide either sparse outcome supervision or richer feedback from process annotations and LLM judges. Outcome rewards scale readily but cannot distinguish grou