返回全部动态

验证而非信任:面向大规模视频发现检索的智能体模型开发

原标题:Verify, Don't Trust: Agentic Model Development for Video Discovery Retrieval at Scale

arXiv cs.IR一手来源研究质量 80

AI 摘要

论文提出 EvoPilot,一种面向长周期在线自动研究的人类把关方法,通过版本化领域技能、类型化适配器和确定性检查来验证实验结果。在 Video Deep Dive 的 37 天检索模型开发中,早期自动研究错误地将 22 个百分点的离线命中率下降归因于交互头,EvoPilot 验证后发现是评估缺陷导致输出深度异常,修复后匹配比较显示离线提升 3.20 个百分点。后续七天随机在线评估显示 VDD 的 GSRR 相对提升 0.66%。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Verify, Don’t Trust: Agentic Model Development for Video Discovery Retrieval at Scale Abstract. Large language model (LLM) agents can now propose, implement, and evaluate model changes. Existing autoresearch loops demonstrate this capability through minutes-scale iterations on a self-contained program. Online autoresearch instead spans asynchronous systems, hours-long variants, and weeks-long campaigns that can influence a product. A completed run can still support an invalid conclusion when a c


发布时间:2026-09-21 12:00
抓取时间:2026-09-21 12:58
来源机构:arXiv
阅读原文arxiv.org