稳定曲线,不稳定项目:视频大语言模型中的项目级缩放异质性
Stable Curves, Unstable Items: Item-Level Scaling Heterogeneity in Video LLMs
AI总结:
该研究发现视频大语言模型存在项目级缩放异质性,单预算无法适配所有项目,通过推导指标验证该效应,提出干预方法并发布审计工件,还展示了响应矩阵的实际应用。
AI中文摘要:
聚合缩放曲线表明,随着视觉预算的增加,视频大语言模型(Video LLMs)会平稳提升或达到饱和状态。我们发现这一观点可能掩盖了项目层面存在的巨大、相反的变化。我们为每个冻结的模型-项目对,通过受控视觉预算下的响应轨迹来表示,并推导得到匹配网格的配置互补性、有害转换和文本覆盖度指标。在来自三个架构家族的五个开源视频大语言模型、四个多项选择题基准拆分、开放式问答与摘要任务,以及固定历史的对话生成任务中,没有单一的预算适用于所有项目。在四模型匹配的多项选择题问答(MCQA)网格上,项目级最优上限跨度为8.8至18.9个准确率点,且12.5%至25.5%的项目在较低预算下正确,但在较高预算下错误。任务适配的连续指标在多项选择题之外也显示出相同的互补性:在MLVU生成任务上,Token-F1最优差距为2.7至3.7个分数点;在AVSD当前轮次生成任务上,该差距为3.8至4.8个点,即使平均质量随预算提升仍存在此效应。该效应在帧率、空间分辨率、采样策略、时空分配,以及独立执行的原始视频和缓存管道中均持续存在,且受项目级比率和成员跟踪协议选择的影响。受控采样干预可恢复29.0%的终端退化,结构化帧审计还识别出几种反复出现的证据路径。我们发布项目级轨迹、协议来源、衍生注释及可复现分析代码作为审计工件。此外,置信度级联方法在匹配固定128帧(128f)准确率的同时,将平均共享帧成本降低了31.7%,展示了响应矩阵的一项实际应用。
英文摘要:
Aggregate scaling curves suggest that Video LLMs improve smoothly or saturate as visual budgets grow. We show that this view can conceal large, opposing changes at the item level. We represent each frozen model--item pair by its response trajectory under controlled visual budgets and derive matched-grid measures of configuration complementarity, harmful transitions, and text overwrite. Across five open Video LLMs from three architecture families, four multiple-choice benchmark splits, open-ended QA and summarization, and fixed-history dialogue generation, no single budget serves all items. On the four-model matched MCQA grid, item-level oracle headroom spans $8.8$--$18.9$ accuracy points and $12.5$--$25.5\%$ of items are correct at a lower budget but wrong at a higher one. Task-appropriate continuous metrics show the same complementarity beyond multiple choice: Token-F1 oracle gaps are $2.7$--$3.7$ score points on MLVU generation and $3.8$--$4.8$ points on AVSD current-turn generation, even when mean quality improves with budget. The effect persists across frame count, spatial resolution, sampling policy, temporal--spatial allocation, and independently executed raw-video and cached pipelines, with per-item rates and membership tracking protocol choices. A controlled sampling intervention recovers $29.0\%$ of terminal regressions, and a structured frame audit identifies several recurring evidence pathways. We release per-item trajectories, protocol provenance, derived annotations, and reproducible analysis code as an auditing artifact. A confidence cascade matches fixed-$128f$ accuracy while reducing average shared frame cost by $31.7\%$, illustrating one operational use of the response matrix.