返回全部动态

闭式与合成孪生:从嵌入统计预测近似最近邻召回率

原标题:Closed Forms and Synthetic Twins: Predicting Approximate Nearest Neighbor Recall from Embedding Statistics

arXiv cs.IR一手来源研究质量 88

AI 摘要

本文提出从嵌入统计量预测近似最近邻索引召回率的方法,包括闭式矩统计、合成孪生语料库模拟和轻校准孪生技术,预测误差在0.03以内。研究还发现可通过训练调整得分边际来提升所有索引族的召回率,从而避免逐语料库的变换。该方法支持索引选择、纠错定价和生产召回预测,适用于持续变化的语料库。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Closed Forms and Synthetic Twins: Predicting Approximate Nearest Neighbor Recall from Embedding Statistics Abstract Embedding models are trained and evaluated as if retrieval were exact; in production they serve behind approximate indexes — HNSW, IVF, product quantization, or the fixed-dimensional encodings (FDEs) of late-interaction models — whose behavior the encoder’s benchmarks never see: one modern encoder recovers just 14% of its exact top-10 through its raw FDE index. Such failures surfac


发布时间:2026-09-02 12:00
抓取时间:2026-09-02 12:15
来源机构:arXiv
阅读原文arxiv.org