澄清后搜索:面向深度搜索的澄清基准与端到端信息恢复
原标题:Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration
AI 摘要
该研究提出了一个用于评估深度搜索中澄清能力的基准,基于百度搜索的真实查询构建了518个实例,并采用闭卷协议和静态金标准来防止意图泄漏。实验表明,澄清能提升搜索效用,其中ERNIE-4.5-Turbo-128K在开源模型中表现最佳,甚至超过GPT-5.2等闭源模型。研究还发现,在严格闭卷条件下,系统常提出无法回答的区域性问题,导致交互效率低下。
正文节选
by Clarify-Then-Search: A Clarification Benchmark for Deep Search with End-to-End Nugget Restoration Abstract. Deep search is brittle on underspecified user queries: missing constraints (e.g., time, location, scope, or definitions) often lead to retrieval drift and incomplete answers. Although LLMs can ask clarifying questions, existing evaluations of clarification for deep search remain limited. We introduce a clarify-then-search benchmark to evaluate the ability of LLMs to ask clarifying quest