返回全部动态

Qdrant 发布 100 亿向量开源基准数据集及 Supernova 框架

原标题:Internet-Scale Knowledge Retrieval: A Novel Vector Search Dataset at 10B Scale

Hugging Face Blog一手来源开源质量 84

AI 摘要

Hugging Face 与 Qdrant、Vultr 合作发布了 Qdrant-FineWeb-10B,这是一个包含 100 亿向量的开源向量搜索基准数据集,基于 FineWeb 语料库生成。同时,Qdrant 开源了 Supernova 框架,用于自动化嵌入生成、暴力 ground truth 计算、数据库加载和评估。该数据集和框架旨在解决现有基准在规模、ground truth 和混合搜索方面的不足,推动大规模向量搜索的标准化评估。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Hugging Face has pioneered open, large-scale model and dataset sharing – setting the bar for data accessibility, findability, and interoperability across the AI community. They currently host some of the largest pre-embedded datasets to date, such as: - CohereLabs/wikipedia-2023-11-embed-multilingual-v3 (~250M vectors, Cohere Embed v3) - Upstash/wikipedia-2024-06-bge-m3 (~144M vectors, BGE-M3) - CohereLabs/msmarco-v2.1-embed-english-v3 (~113M vectors, Cohere Embed v3) - Qdrant/dbpedia-entities-o


发布时间:—
抓取时间:2026-09-07 07:47
来源机构:Hugging Face
阅读原文huggingface.co