arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

参数化多模态用户记忆:存储字幕无法承载的内容

Parametric Multimodal User Memory: Storing What Captions Cannot Carry

Bojie Li, Noah Shi

arXiv 2608.28609首次发表:更新:

发表机构

Pine AI; University of Washington(派恩人工智能; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出参数化多模态用户记忆,结合视觉语言模型与专用编码器,解决传统文本用户记忆丢失感知信息的问题,在PerceptMem数据集上验证了其有效性与可推广性。

AI 中文摘要

个性化智能体需要用户记忆,即关于用户是谁的持久模型。如今,这种记忆几乎总是文本形式——通过相似度检索的转录文本和字幕,这能覆盖用户可被字幕描述的那部分(“我的猫叫比比”),但会丢失字幕无法承载的感知部分:声音的音色、不同年龄和光照下的面部特征、某人说话时的疲惫程度等。我们在五种模态上测量了这种损失:基于强字幕的重识别模型的召回率仅为专用编码器的0.11,在不可命名信号上会坍缩至随机水平。我们转而将感知记忆建立在模型内部,将召回分解为两个子问题:视觉语言模型(VLM)在上下文中标定指代对象(是什么、在哪里),专用编码器提取身份密钥(是谁),存储为生成时注意力可读取的内联令牌,无需外部往返。两者单独使用均不足够:VLM跨年龄面部识别召回率仅为0.54,而面部编码器达到0.81;未标定的编码器对两人场景指代对象的识别率为0.05;但两者结合后达到正确区域的神谕水平(0.96),并可推广至多说话者音频和视频。该识别核心无需训练:它在O(1)配准成本下,于任何冻结模型上复现编码器的召回率。在PerceptMem(12个领域、1080个任务)中,感知身份受容量限制,而精确事实受绑定限制:身份应属于参数化库,事实属于文本存储。两种记忆可干净组合:拥有两者的智能体不仅能记住用户说过的话,还能记住用户的特质。

英文摘要

A personalized agent needs a user memory: a persistent model of who its user is. Today it is almost always text -- transcripts and captions retrieved by similarity. This serves the captionable half of a person ("my cat is named Bibi"), but discards the perceptual half no caption can hold: how a voice sounds, how a face reads across age and lighting, how tired someone sounds. We measure this loss across five modalities: a strong caption-based re-identifier recovers as little as 0.11 of a dedicated encoder's recall, collapsing toward chance on non-nameable signals. We instead ground perceptual memory in the model, decomposing recall into two subproblems: a vision-language model grounds the referent in context (what and where), and a dedicated encoder extracts an identity key (who), stored as one inline token read by attention at generation with no external round-trip. Neither suffices alone -- the VLM identifies cross-age faces at only 0.54 recall where a face encoder reaches 0.81, and an ungrounded encoder recognizes a two-person-scene referent at 0.05 -- yet together they reach correct-region oracle (0.96), generalizing to multi-speaker audio and video. The recognition core is training-free: it reproduces the encoder's recall on any frozen model at O(1) registration cost. On PerceptMem (12 domains, 1,080 tasks) perceptual identity is capacity-limited while exact facts are binding-limited: identity belongs in a parametric bank, facts in a text store. The two memories compose cleanly: an agent with both can remember not only what its user said, but also what they are like.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑