返回全部动态

通过轨迹微调让小型语言模型学习检索动作

原标题:The Fellowship of the Query: Learning Retrieval Actions

arXiv cs.IR一手来源研究质量 80

AI 摘要

该论文研究通过轨迹微调提升小语言模型(SLM)作为检索增强问答中「下一步动作控制器」的能力。作者从教师模型 Qwen 3-Next 80B-A3B Instruct 的搜索轨迹中构建七类动作预测任务,并用 LoRA 监督微调多个 SLM 与 xSLM。在 1,646 个留出动作样本上,Granite 4.1 3B 的 macro-F1 达 0.6536,远高于零样本提示的 0.1736 和 TF-IDF 逻辑回归基线的 0.5399;端到端评估中 Exact Match 从 0.7530 提升至 0.7946,token F1 从 0.7783 提升至 0.8295。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

The Fellowship of the Query: Learning Retrieval Actions Abstract. Retrieval-augmented question answering requires control decisions about when to decompose a question, search, reformulate, extract evidence, synthesize facts, verify progress, and stop. We study whether trajectory fine-tuning can improve small language models (SLMs) as next-action controllers. We additionally evaluate a low-resource setting in which a single SLM serves as both the controller and the final-answer generator. From ac


发布时间:2026-09-25 12:00
抓取时间:2026-09-25 12:10
来源机构:arXiv
阅读原文arxiv.org