arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.29607cs.CVcs.CL

STRAND:视频大语言模型中以对象为中心的空间-时间监控的基准测试与改进

STRAND: Benchmarking and Improving Object-Centric Spatio-Temporal Monitoring in Video Large Language Models

Thong Nguyen, Tri Cao, Khoi Le, Cong-Duy Nguyen, Quynh Vo, See-Kiong Ng, Bryan Hooi Kuen-Yew

首次发表
浏览论文内容

中文总结 AI 辅助

针对视频大语言模型在动态场景中的幻觉问题,提出STRAND基准以子问题分解和忠实准确率评估时空监控能力,并构建以对象为中心的轨迹推理框架,显著减少幻觉并提升推理一致性。

中文摘要 AI 辅助

尽管多模态大语言模型(MLLMs)在视频理解方面取得了进展,但在动态场景中仍然极易产生幻觉。我们认为这源于空间-时间监控的失败,即持续跟踪对象身份、状态和关系的能力。现有基准通过依赖对通常可以通过局部视觉线索或统计先验解决的查询进行单一最终答案评估,掩盖了这一缺陷。为了严格诊断这一问题,我们引入了STRAND,一个经过人工验证的以对象为中心的事实的基准,通过将查询分解为子问题来评估中间推理,区分真正的时序理解与偶然的正确性。关键在于,我们使用忠实准确率(Faithful Accuracy)对模型进行评分,这是一种无条件的联合指标,仅当目标答案和每个前提子问题都正确时才给予预测信任,因此模型无法通过在恰好正确回答的少量目标子集上选择性保持一致来夸大其分数。为了解决STRAND暴露的失败模式,我们进一步提出了一个以对象为中心的框架,通过分块状态提取和时序聚合显式构建并推理结构化对象轨迹。大量实验,包括与端到端MLLMs和模块化视频框架的骨干、帧、调用和令牌匹配比较,表明我们的以对象为中心的框架显著减少了幻觉答案,并提高了相对于最先进MLLMs的空间-时间推理一致性。代码、模型和数据已在此HTTP URL上提供。

英文摘要

While multimodal large language models (MLLMs) have advanced video understanding, they remain highly prone to hallucinations in dynamic scenes. We argue this stems from a failure in spatio-temporal monitoring, the ability to persistently track object identities, states, and relations over time. Existing benchmarks obscure this deficit by relying on single final-answer evaluations for queries that can often be resolved via local visual cues or statistical priors. To rigorously diagnose this, we introduce STRAND, a benchmark of human-verified object-centric facts that evaluates intermediate reasoning by decomposing queries into sub-questions, distinguishing genuine temporal understanding from coincidental correctness. Crucially, we score models with Faithful Accuracy, an unconditional joint metric that credits a prediction only when the target answer and every prerequisite sub-question are correct, so that a model cannot inflate its score by being selectively consistent on the small subset of targets it happens to answer correctly. To address failure modes exposed by STRAND, we further propose an object-centric framework that explicitly constructs and reasons over structured object trajectories via chunk-wise state extraction and temporal aggregation. Extensive experiments, including backbone-, frame-, call-, and token-matched comparisons against both end-to-end MLLMs and modular video harnesses, demonstrate that our object-centric framework significantly reduces hallucinated answers and improves spatio-temporal reasoning consistency over state-of-the-art MLLMs. The code, model, and data have been made available at nguyentthong.github.io/strand.

发表机构

  • National University of Singapore (NUS)(新加坡国立大学)
  • Centre for AI Research, VinUniversity(VinUniversity人工智能研究中心)

机构由 AI 辅助整理,请以论文原文为准。

↑