Qdrant 发布 100 亿向量开源基准数据集及 Supernova 框架
原标题:Internet-Scale Knowledge Retrieval: A Novel Vector Search Dataset at 10B Scale
AI 摘要
Hugging Face 与 Qdrant、Vultr 合作发布了 Qdrant-FineWeb-10B,这是一个包含 100 亿向量的开源向量搜索基准数据集,基于 FineWeb 语料库生成。同时,Qdrant 开源了 Supernova 框架,用于自动化嵌入生成、暴力 ground truth 计算、数据库加载和评估。该数据集和框架旨在解决现有基准在规模、ground truth 和混合搜索方面的不足,推动大规模向量搜索的标准化评估。
正文节选
Hugging Face has pioneered open, large-scale model and dataset sharing – setting the bar for data accessibility, findability, and interoperability across the AI community. They currently host some of the largest pre-embedded datasets to date, such as: - CohereLabs/wikipedia-2023-11-embed-multilingual-v3 (~250M vectors, Cohere Embed v3) - Upstash/wikipedia-2024-06-bge-m3 (~144M vectors, BGE-M3) - CohereLabs/msmarco-v2.1-embed-english-v3 (~113M vectors, Cohere Embed v3) - Qdrant/dbpedia-entities-o