发表机构
KAIST; VGG, University of Oxford(韩国科学技术院; 牛津大学视觉几何组)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究旨在降低Omni-LLMs推理时输入令牌成本,提出无需训练的ReMo框架,通过跨模态重分配信息压缩视觉令牌,在Qwen2.5-Omni上实验,去除54%输入令牌且准确率无损失,甚至略超全令牌模型。
AI 中文摘要
本文旨在降低全模态大语言模型(Omni-LLMs)推理时的输入令牌成本。Omni-LLMs能跨音频、视频和文本联合推理,但三模态流成本高度不平衡,视觉令牌占输入大部分且冗余。我们提出ReMo,一个无需训练的框架,通过跨模态重新分配信息来压缩视觉令牌:仅当视觉令牌信息在其他地方未出现时才保留。ReMo通过两种方式实现:一是在公共嵌入空间中对齐音频和视频,去除已由音频或其他视觉令牌解释的视觉令牌;二是用紧凑文本代理替换对象级视觉令牌。在两个模型规模的Qwen2.5-Omni上,ReMo去除54%的输入令牌且准确率无损失,甚至略超全令牌模型,在五个视听基准上达到其平均准确率的101.2%和101.3%。
英文摘要
The goal of this paper is to reduce the input token cost of Omni-modal large language models (Omni-LLMs) at inference time. Omni-LLMs reason jointly over audio, video and text, but the cost of the three streams is highly unbalanced: visual tokens account for the vast majority of the input, and are highly redundant. In this paper, we propose ReMo, a training-free framework that compresses visual tokens by redistributing their information across modalities: a visual token is kept only if its information appears nowhere else. ReMo achieves this in two ways: (i) it aligns audio and video in a common embedding space, and removes visual tokens already explained by the audio or by other visual tokens; and (ii) it replaces object-level visual tokens with compact text proxies, short descriptions of each object and its location, conveying the same content in far fewer tokens. On Qwen2.5-Omni at two model scales, ReMo removes 54% of the input tokens with no loss in accuracy. Indeed, it slightly exceeds the full-token model, reaching 101.2% and 101.3% of its average accuracy over five audio-visual benchmarks.
CommentsPreprint