arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

查询无关的可变码率视觉标记编码

Query Independent Variable Rate Visual Token Coding

Hongbo Zhang, Zihao Yang, Liuyang Song, Daqian Yang, Haoyang Yao, Yan Wen, Zhengtao Yao

arXiv 2610.00204首次发表:更新:

发表机构

Peking University; Uppsala University; University of Southern California(北京大学; 乌普萨拉大学; 南加州大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对视觉-语言模型,提出查询无关的可变码率视觉标记编码,通过失真-码率曲线分配比特预算,在保持输出分布和答案上优于均匀码率及剪枝方法,并超越注意力排序剪枝于跨问题场景。

AI 中文摘要

视觉-语言模型的视觉标记压缩几乎完全被设定为一个选择问题:决定保留哪些标记并丢弃其余部分。最有效的标准是根据语言模型对标记的关注度进行排序,这使得排序成为所提问题的函数。这在单轮基准测试中不可见,但在压缩表示被写入一次并多次读取时(例如在对话轮次之间缓存或在设备与服务器之间传输)则具有决定性作用。我们转而采用经典变换编码工具包的另一半:保留每个标记并改变其码率。变换编码暴露每个标记的实测失真-码率曲线,并通过在这些曲线上进行精确的整数码率-失真优化,将固定比特预算分配给各标记。文本不进入流程,因此一个压缩表示可服务于任何查询。在相同比特预算下,在两个数据集和两种容量上,它比均匀码率编码、闭式注水法和失真排序剪枝更好地保留了模型的输出分布及其答案。它在问题剪枝所针对的问题上匹配了注意力排序剪枝,并且一旦压缩图像必须回答关于同一图像的不同问题时,它便超越了注意力排序剪枝。

英文摘要

Visual-token compression for vision--language models is posed almost entirely as a selection problem: decide which tokens to keep and discard the rest. The criteria that work best rank tokens by the attention the language model pays them, which makes the ranking a function of the question being asked. That is invisible in a single-turn benchmark and decisive whenever a compressed representation is written once and read many times, as when it is cached across the turns of a conversation or transmitted between a device and a server. We take the other half of the classical transform-coding toolkit instead: keep every token and vary its rate. A transform code exposes each token's measured distortion--rate curve, and a fixed bit budget is distributed across tokens by exact integer rate--distortion optimisation on those curves. No text enters the pipeline, so one compressed representation serves any query. At equal bit budgets, on two datasets and two capacities, it preserves the model's output distribution and its answers better than uniform-rate coding, the closed-form water-fill and distortion-ranked pruning. It matches attention-ranked pruning on the question pruning was tuned for, and overtakes it once the compressed image must answer a different question about the same image.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑