Google DeepMind联合Kaggle发布FACTS基准套件,系统评估大模型事实性
原标题:FACTS Benchmark Suite: Systematically evaluating the factuality of large language models
AI 摘要
Google DeepMind与Kaggle合作推出FACTS Benchmark Suite,扩展了原有FACTS Grounding Benchmark,新增参数化、搜索和多模态三个事实性基准,并更新了Grounding Benchmark v2。该套件共包含3,513个公开示例,由Kaggle管理私有测试集并主持公共排行榜,用于系统评估大语言模型的事实准确性。
正文节选
Large language models (LLMs) are increasingly becoming a primary source for information delivery across diverse use cases, so it’s important that their responses are factually accurate. In order to continue improving their performance on this industry-wide challenge, we have to better understand the types of use cases where models struggle to provide an accurate response and better measure factuality performance in those areas. The FACTS Benchmark Suite Today, we’re teaming up with Kaggle to int