返回全部动态

SPIEval:评估大语言模型作为移动助手处理分散个人信息的能力

原标题:SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information

Hugging Face Daily Papers一手来源研究质量 82

AI 摘要

SPIEval是一个新基准,用于评估大语言模型作为移动助手处理分散在多个应用中的个人信息的能力。该基准包含250个任务,覆盖10个应用中的4335条个人记录,并支持通过21个工具进行多轮交互。评估发现,最佳模型GPT-5.5(xhigh)准确率仅为57.3%,最弱模型为16.4%,且79%的失败源于信息定位不准确。研究揭示了当前基于LLM的移动助手的根本局限性,并提供了数据和代码。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

SPIEval: Evaluating Large Language Models as Mobile Assistants over Scattered Personal Information Abstract SPIEval benchmarks mobile assistant LLMs on scattered personal data tasks, revealing major gaps in information retrieval and verification. Large language models (LLMs) are increasingly deployed as mobile assistants, where a key challenge is leveraging personal information scattered across multiple applications (apps) to complete user instructions. However, due to the lack of dedicated benc


发布时间:—
抓取时间:2026-08-12 11:02
来源机构:Hugging Face
阅读原文huggingface.co