返回全部动态

参数化多模态用户记忆:存储字幕无法承载的感知信息

原标题:Parametric Multimodal User Memory: Storing What Captions Cannot Carry

arXiv cs.CL一手来源研究质量 87

AI 摘要

本文提出了一种参数化多模态用户记忆方法,将感知记忆(如声音、面部)以原生模态存储,通过VLM进行指代定位、专用编码器提取身份键,并利用AttMem记忆库以注意力令牌形式存储,无需外部检索。实验表明,该方法在跨年龄、多说话人等场景下显著优于基于文本描述的记忆系统,弥补了纯文本记忆的不足。

以上摘要由 AI 生成,可能存在误差。事实请以原文为准。

正文节选

Parametric Multimodal User Memory: Storing What Captions Cannot Carry Abstract A personalized agent needs a user memory: a persistent model of who its user is. Today it is almost always text — transcripts and captions retrieved by similarity. This serves the captionable half of a person (“my cat is named Bibi”), but discards the perceptual half no caption can hold: how a voice sounds, how a face reads across age and lighting, how tired someone sounds. We measure this loss across five modalities:


发布时间:2026-09-01 12:00
抓取时间:2026-09-01 12:02
来源机构:arXiv
阅读原文arxiv.org