返回全部动态

个性化幻象:LLM捏造用户画像,自我监控误导

原标题:The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads

Hugging Face Daily Papers一手来源研究质量 87

AI 摘要

Hugging Face 每日论文发布了一项关于个性化LLM的研究,揭示了过度推断现象:模型会捏造用户属性。研究团队推出了MirageBench基准,包含150个角色和6个任务,评估了12个模型,发现所有模型都存在过度推断,平均41.6%的声明为捏造。最引人注目的是“自我监控反转”:模型自我评估的过度推断水平与外部评判结果呈负相关,表明自我报告不可靠,外部验证更可靠。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

The Personalization Mirage: How LLMs Fabricate User Profiles, and Why Self-Monitoring Misleads Abstract Personalized LLMs with persistent memory are increasingly deployed, yet the faithfulness of their user models remains unexamined. We study over-inference (OI): the phenomenon where LLMs fabricate user attributes beyond what evidence supports. We introduce MirageBench, comprising 150 personas balanced across stereotypical, counter-stereotypical, and neutral profiles, 6 personalization tasks spa


发布时间:—
抓取时间:2026-08-06 20:57
来源机构:Hugging Face
阅读原文huggingface.co