发表机构
The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出PixelTriage,一个置于检索后的小模型,通过读取对话、备注和缩略图预测每条记忆像素的价值,仅用少量视觉标记即可保持准确率并加速多模态助手回答。
AI 中文摘要
多模态助手从包含图像的长时记忆中回答问题。在检索之后,每个检索到的图像要么以像素形式到达回答模型,每张图像约消耗一千个视觉标记,要么以存储的文本代理形式到达,但文本代理常常遗漏问题所询问的细节。我们发现,像素带来的好处通常来自一到两条检索到的记忆,并且可以在回答模型运行之前、无需读取任何全分辨率图像的情况下预测这种好处。在PixelTriage中,这是一个放置在检索之后的插件,一个不生成文本的小型模型读取对话、简短说明以及每条检索记忆的缩略图,并预测其像素会增加多少价值。该模型在由冻结的27B模型标注的合成记忆片段上进行训练,该27B模型在有无每条记忆像素的情况下分别回答每个问题。使用7B回答模型时,PixelTriage位于M$^3$Exam、DMV和MemEye的准确率-成本前沿上,并且仅使用11%至23%的视觉标记,而准确率没有显著损失。在DMV上,它的回答速度比打开所有图像快2.9倍。在同等预算下,它优于检索顺序和均匀缩小尺寸,并可迁移到其他记忆系统和397B回答模型。
英文摘要
Multimodal assistants answer questions from long-term memories that contain images. After retrieval, each retrieved image reaches the answering model either as pixels, at about a thousand visual tokens per image, or as a stored text proxy that often misses the detail the question asks about. We find that the benefit of pixels usually comes from one or two retrieved memories, and that it can be predicted before the answering model runs, without reading any full-resolution image. In PixelTriage, a plug-in placed after retrieval, a small model that does not generate text reads the dialogue, a short note and a thumbnail of each retrieved memory and predicts how much its pixels would add. It is trained on synthetic memory episodes labeled by a frozen 27B model that answers each question with and without each memory's pixels. With a 7B answering model, PixelTriage lies on the accuracy--cost frontier of M$^3$Exam, DMV and MemEye and uses 11--23\% of the visual tokens without a significant loss of accuracy. On DMV it answers 2.9 times faster than opening all images. It outperforms retrieval order and uniform down-sizing at equal budgets and transfers to other memory systems and to a 397B answering model.