返回全部动态

Google DeepMind联合Kaggle发布FACTS基准套件,系统评估大模型事实性

原标题:FACTS Benchmark Suite: Systematically evaluating the factuality of large language models

Google DeepMind News一手来源研究质量 82

AI 摘要

Google DeepMind与Kaggle合作推出FACTS Benchmark Suite,扩展了原有FACTS Grounding Benchmark,新增参数化、搜索和多模态三个事实性基准,并更新了Grounding Benchmark v2。该套件共包含3,513个公开示例,由Kaggle管理私有测试集并主持公共排行榜,用于系统评估大语言模型的事实准确性。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Large language models (LLMs) are increasingly becoming a primary source for information delivery across diverse use cases, so it’s important that their responses are factually accurate. In order to continue improving their performance on this industry-wide challenge, we have to better understand the types of use cases where models struggle to provide an accurate response and better measure factuality performance in those areas. The FACTS Benchmark Suite Today, we’re teaming up with Kaggle to int


发布时间:2025-12-09 19:29
抓取时间:2026-09-07 03:44
来源机构:Google DeepMind
阅读原文deepmind.google