NVIDIA 适配 Nemotron 检索栈并发布希腊语 RAG 基准 HERA
原标题:Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains
AI 摘要
NVIDIA 的研究团队针对现代希腊语对 Nemotron 检索栈进行了端到端适配,包括语料挖掘、合成监督、检索模型训练、重排序器适配和阅读器微调,并发布了名为 HERA 的基准。研究发现 BM25 基线在专业希腊语语料上优于多个现成的多语言稠密检索模型,而微调后的 Nemotron 1B 嵌入器将 nDCG@10 从 0.362 提升至 0.835。此外,通过 LoRA 微调 Nemotron 30B-A3B 混合专家阅读器,答案正确率从 29.4% 提升至 66.9%,同时显著改善了忠实度和引用质量。模型和基准已公开发布,以支持希腊语 RAG 系统的研究。
正文节选
Teaching Nemotron Greek: Mining a Corpus, Adapting Retrieval, and Grounding Generation for Modern Greek across Specialist Domains Abstract Modern Greek is absent from NVIDIA's Nemotron retrieval models and from major multilingual retrieval benchmarks, despite being important for retrieval-augmented generation (RAG) in legal, energy, financial, and medical applications. We present an end-to-end adaptation of the Nemotron retrieval stack for Modern Greek, including corpus mining, synthetic supervi