arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

为过多 Token 付费?使用简单启发式方法进行有效且成本高效的多模态大语言模型标注

Paying for Too Many Tokens? Valid and Cost-Efficient Multimodal LLM Annotation with Simple Heuristics

Zhixi Zhu, Kristina Gligoric

arXiv 2610.00809首次发表:更新:

发表机构

Johns Hopkins University(约翰霍普金斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究系统评估了多模态大语言模型在短视频标注中的多种成本削减启发式方法,发现准确性与结论有效性可能背离,且单模态文本可高效,并证明基于镜头检测的2×8图像网格能以15%的Token成本接近全视频理解性能,为成本感知标注提供指导。

AI 中文摘要

视觉-语言模型(VLMs)能够实现大规模视频标注,但成本会迅速累积:以每秒一帧的速度处理一个典型的60秒短视频需要数百万个Token。为降低成本,研究人员依赖启发式方法,如对帧进行子采样、将视频压缩成图像网格,或仅使用单一模态。然而,目前尚不清楚哪些启发式方法能节省成本,以及它们是否保留了这些标注所支持的下游结论。为弥补这一空白,我们使用短视频,在两个计算社会科学(CSS)任务(情感分类和主题分类)上对这些启发式方法进行了系统评估。我们沿着文献中通常分开处理的三个轴评估每种配置:分类准确性、下游推断的有效性以及每视频的Token成本。首先,我们发现准确性和有效性存在分歧:最高准确性的配置可能产生错误的结论。其次,模态的价值并非有保证:仅文本就能产生强大的性能,这表明添加模态可能会增加成本而不增加信号。最后,我们发现,在标注短视频时,成本可以与视频长度解耦:通过简单的镜头转换检测构建的单个$2\times8$图像网格,其全视频理解性能($\kappa$在0.05以内)接近,而Token成本约为全视频的15%。基于这些发现,我们得出了能够在CSS中实现成本感知的VLM标注的指导原则。

英文摘要

Vision-Language Models (VLMs) enable video annotation at scale, but costs accumulate quickly: processing a typical 60-second short-form video at one frame per second requires millions of tokens. To reduce costs, researchers rely on heuristics such as sampling a subset of frames, compressing videos into image grids, or using only a single modality. However, it remains unclear which heuristics save cost, and whether they preserve the downstream conclusions these annotations enable. To address this gap, we conduct a systematic evaluation of these heuristics using short-form videos, on two computational social science (CSS) tasks: sentiment and topic classification. We evaluate each configuration along three axes the literature typically treats separately: classification accuracy, validity of downstream inference, and per-video token cost. First, we find that accuracy and validity diverge: the highest-accuracy configuration can produce wrong conclusions. Second, modality value is not guaranteed: text alone can yield strong performance, indicating that adding modalities can add cost without adding signal. Finally, we find that cost can be decoupled from video length when annotating short-form videos: a single $2\times8$ image grid built via simple shot-transition detection approaches full-video understanding ($κ$ within~.05), at $\sim 15\%$ of the token cost. Based on these findings, we derive guidelines that can enable cost-aware VLM annotation in CSS.

Journal refAACL 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑