arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30670cs.CLcs.CV

TRACE:流式视频理解的时间审计与条件感知评估

TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

  • Om AI Research(Om AI研究院)

机构由 AI 辅助整理,请以论文原文为准。

Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao

AI总结:

针对流式视频理解评估忽视证据时序与触发条件的问题,提出TRACE基准框架,通过条件感知协议和多维指标揭示准确率相同背后工作负载与行为差异。

AI中文摘要:

流式视频理解要求模型在证据到达时进行解释,然而当前的评估往往仅报告任务得分,而未指明证据何时变得有效、视觉历史如何维护或响应如何触发。因此,相似的得分可能对应不同的工作负载、失败模式和操作行为。我们提出了TRACE(时间审计与条件感知评估),一个使这些因素明确化的条件感知基准和评估框架。TRACE结合了带证据时间和指令相关触发注释的时间审计视觉任务、一个统一的因果核心-适配器协议(该协议在记录实际历史处理和响应事件的同时控制信息可用性),以及关于答案质量、及时性、响应选择行为、工作负载、完成度和可靠性的多维报告。在来自517个视频的1240条记录上,我们评估了八种配置下的八个公开可用的模型或系统。我们发现,几乎相同的问答准确率可能掩盖完成度、答案有效性和生成工作负载方面的显著差异,而主动性能则区分为响应质量、响应延迟、误报(在没有目标窗口当前有效且稍后仍存在时发出的响应)和错过的目标窗口。这些结果表明,流式视频性能应被解释为执行条件下的系统行为,而非单一得分。我们的基准和代码可在以下网址访问:此https URL。

英文摘要:

Streaming video understanding requires models to interpret evidence as it arrives, yet current evaluations often report task scores without specifying when evidence becomes valid, how visual history is maintained, or how responses are triggered. As a result, similar scores may correspond to different workloads, failure modes, and operational behavior. We introduce TRACE (Temporal Audit and Condition-aware Evaluation), a condition-aware benchmark and evaluation framework that makes these factors explicit. TRACE combines temporally audited visual tasks with evidence timing and instruction-dependent trigger annotations, a unified causal Core--Adapter protocol that controls information availability while recording actual history processing and response events, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. On 1,240 records from 517 videos, we evaluate eight publicly available models or systems in eight configurations. We find that nearly identical QA accuracy can mask substantial differences in completion, answer validity, and generation workload, while proactive performance separates into response quality, response delay, false alarms (responses emitted while no target window is currently valid and a later one remains), and missed target windows. These results show that streaming-video performance should be interpreted as execution-conditioned system behavior rather than a single score. Our benchmark and code can be accessed at \href{https://github.com/om-ai-lab/trace-bench}{https://github.com/om-ai-lab/trace-bench}.

补充信息

↑