选择、压缩、再投入:长视频多模态大语言模型中视觉令牌分配的受控研究
Select, Compress, Reinvest: A Controlled Study of Visual-Token Allocation in Long-Video MLLMs
浏览论文内容
中文总结 AI 辅助
本文通过控制变量的受控研究,发现长视频多模态大语言模型中视觉令牌分配的选择策略影响最大,正交匹配追踪算法性能接近专用选择器,压缩几乎无成本,再投入可提升精度。
中文摘要 AI 辅助
长视频语言模型无法查看每一帧:以每秒采样一次的时长为一小时的视频,共有3600张图像,而系统仅保留其中一小部分固定切片。该切片保留哪些帧通常被视为预处理细节;本文验证了这是否重要。已发表的选择器因同时修改帧评分器、提示边界、分辨率策略和回答模型,导致难以进行比较。本文控制其余变量固定,仅逐一改变三个决策:选择、空间压缩和节省令牌的再投入,覆盖6种无训练选择规则、3个长视频基准和2个回答模型。选择是最大的单一影响因素:在LongVideoBench的小时级数据集中,8个经查询选择的帧比16个均匀间隔的帧高出6.9个点;正交匹配追踪(Orthogonal Matching Pursuit)这一已有数十年历史的未修改稀疏近似算法,在所有三个基准中,其性能与所有专用选择器相当或差距在1个点以内。压缩几乎无成本:在固定时间戳将每帧空间预算减半,最多仅损失0.44个点。再投入是将节省的预算转化为精度的关键:将释放的令牌用于两倍数量的压缩帧,在不超过原始8个帧的测量成本下,可进一步提升2至3个点;压缩只有在以这种方式使用节省的令牌时才会产生收益。研究过程中,本文发现自身AKS基线存在一个实现错误,以及同一已发表规则在相同预算下,两个测试框架间存在0.07至3.74个点的差距,这表明此类比较需在同一受控框架内进行,而非跨论文开展。
英文摘要
Long-video language models cannot look at every frame: an hour sampled once per second is 3,600 images, and a system keeps only a small fixed slice of that pool. Which frames survive that slice is usually treated as a preprocessing detail; we test whether it should be. Published selectors make the comparison hard because they change the frame scorer, the prompt boundary, the resolution policy, and the answering model all at once. We hold each fixed and vary one decision at a time: selection, spatial compression, and reinvestment of the savings, across six training-free selection rules, three long-video benchmarks, and two answering models. Selection is the largest single lever: on LongVideoBench's hour-long bin, eight query-selected frames beat sixteen uniformly spaced ones by 6.9 points, and Orthogonal Matching Pursuit, an unmodified decades-old sparse-approximation algorithm, matches or comes within a point of every purpose-built selector we compare it against, across all three benchmarks. Compression is close to free: halving each frame's spatial budget at fixed timestamps costs at most 0.44 points. Reinvestment is where that budget turns back into accuracy: spending the freed tokens on twice as many compressed frames, at a measured cost no higher than the original eight, returns a further two to three points; compression only pays off once its savings are spent this way. Along the way, an implementation bug in our own AKS baseline and a 0.07 to 3.74 point gap between two harnesses running the same published rules at the same budget show why these comparisons need to happen inside one controlled harness rather than across papers.