返回全部动态

LightOn 发布开源多语言检索模型 mDenseOn 和 mLateOn

原标题:mDenseOn with the mLateOn: Open Multilingual, Long-Context, and Code Retrieval Models

Hugging Face Blog一手来源模型发布质量 82

AI 摘要

LightOn AI 发布了两个开源的 307M 参数多语言检索模型 mDenseOn 和 mLateOn,基于 28 亿对翻译训练语料,覆盖 8 种语言和代码。mLateOn 在 BEIR、MIRACL 和 MLDR 基准上表现最佳,尤其在未见语言上泛化能力显著优于 mDenseOn。所有模型、数据集和训练代码均已开源。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

We are releasing mDenseOn and mLateOn, two open 307M-parameter multilingual retrieval models trained on a 2.8B-pair translate-train corpus, one of the largest open multilingual (and cross-lingual) retrieval training sets to date. They extend the open English data recipe we validated with DenseOn and LateOn to eight additional languages. In our evaluation, mLateOn gets the best results on BEIR, on target-language MIRACL, and on both target-language and full MLDR, and it stays competitive on full


发布时间:—
抓取时间:2026-08-02 00:09
来源机构:Hugging Face
阅读原文huggingface.co