返回全部动态
检索需要多向量:单向量与多向量嵌入的指数级分离
原标题:Retrieval Needs Multivectors: An Exponential Separation
AI 摘要
该研究首次证明了在检索排序任务中,单向量嵌入与多向量嵌入之间存在指数级表达能力差距。作者构造了显式的查询-文档相关性矩阵族,其中多向量嵌入只需多项式大小即可正确排序,而单向量嵌入需要指数大小。基于理论构造,他们提出了ANDOR基准,实验显示单向量模型在零样本和微调后均表现不佳,而多向量模型持续领先,验证了理论预测。
以上摘要由 AI 生成,可能存在误差。事实请以原文为准。
正文节选
Retrieval Needs Multivectors: An Exponential Separation Abstract Recent works have highlighted the expressive limitations of embedding based retrieval models through both theoretical analyses and challenging benchmarks such as LIMIT. While multi-vector embeddings consistently outperform single-vector embeddings, the precise representational gap between them remains poorly understood. In this work, following Jayaram’s work, we provide the first explicit family of query and document sets, together
发布时间:2026-08-25 12:00
抓取时间:2026-08-25 12:50
来源机构:arXiv