arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

屈服还是信服:时序采样门控视频大语言模型的主张依从性

Caved or Convinced: Temporal Sampling Gates Claim Deference in Video Large Language Models

Yuxin Cao, Wei Song, Jingling Xue, Jin Song Dong

arXiv 2608.03160首次发表:更新:

发表机构

National University of Singapore; University of New South Wales(新加坡国立大学; 新南威尔士大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对视频大语言模型的两种失败情况,区分了可用性与权重两个原因,提出反转测试缓解屈服于错误主张的问题,提升了顺序准确率并使模型可弃权而非猜测。

AI 中文摘要

当被问及两个事件中哪个先发生时,视频大语言模型会出现两种相反的失败情况:屈服于错误主张,或拒绝真实主张。先前关于视频谄媚的研究仅衡量前者,并通过教导模型减少对用户的信任来缓解该问题,而该修复方法在文本和图像模型中会恶化后者。在视频领域,这两种失败源于文献视为同一的两个原因:可用性(即稀疏采样的帧是否包含这两个事件)和权重(即该证据是否被优先于用户的主张)。我们通过两项干预措施将二者区分,同时保持主张固定:一种保留帧的重排(可翻转主张的真实性),以及一种采样偏移(在固定帧预算下捕获或遗漏两个事件)。当事件被遗漏时,两个孪生样本呈现相同帧,因此我们评估的9个模型中每个都以相同速率接受真实和错误主张,这使得Youden's J在构造上为0。可用性是必要条件但非充分条件。9个模型中有5个能读取顺序,但其中4个仍屈服于错误主张,因此它们的依从性达到权重上限。由于信任无法在从未采样的证据上校准,我们提出一种反转测试,通过对采样帧正向和反向评分来抵消模型的顺序先验,随后在不读取主张的情况下回答、重采样或弃权(不执行)。该测试将能读取顺序的模型的顺序准确率提升至0.92-1.00,并在无法读取顺序的模型上使其弃权而非猜测。

英文摘要

When asked which of two events came first, video large language models can fail in two opposite ways: cave to a false claim, or reject a true one. Prior video sycophancy work measures only the first and mitigates it by teaching the model to trust the user less, a fix known in text and image models to worsen the second. In video, both failures come from two causes the literature treats as one: availability, whether the sparse sampled frames contain the two events, and weighting, whether that evidence is trusted over the user. We separate them with two interventions that keep the claim fixed: a frame-preserving reorder that flips the claim's truth, and a sampling-offset shift that captures or misses both events at a fixed frame budget. When the events are missed, the two twins present identical frames, so each of the nine models we evaluate accepts a true and a false claim at the same rate, making Youden's $J=0$ by construction. Availability is necessary but not sufficient. Five of the nine read the order, yet four of those five still cave to the false claim, so their deference hits a weighting ceiling. Since trust cannot be calibrated over evidence that was never sampled, we propose a reversal test that cancels the model's order prior by scoring the sampled frames forward and reversed, then answers, resamples, or abstains without reading the claim. The test raises the order accuracy to 0.92-1.00 on the models that read the order and abstains rather than guesses on those that cannot.

Comments11 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑