arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.30722cs.CVcs.AI

TrafficImag:用于反事实路侧交通视频生成的基准

TrafficImag: A Benchmark for Counterfactual Roadside Traffic Video Generation

Xiangyu Li, Tianyi Wang, Zhihao Dou, Christian Claudel, Zhaomiao Guo

首次发表
浏览论文内容

中文总结 AI 辅助

TrafficImag是首个反事实路侧交通视频生成基准,结合大规模数据集与可执行协议,评估四个有效性维度,发现条件视频执行是主要瓶颈。

中文摘要 AI 辅助

现有的路侧交通数据集支持感知、预测和视觉问答,但它们不评估反事实视频生成,即在反事实视频生成中,选定的参与者被修改,生成的未来应与道路拓扑和无关交通保持一致。我们引入了TrafficImag,这是第一个用于反事实路侧交通视频生成的基准。TrafficImag结合了一个大规模路侧数据集(9,022张标注图像、7,043个去重视频片段和31,145个以参与者为中心的历史-未来样本)与一个可执行的协议,该协议支持行为推理、干预感知的图像编辑和条件视频生成。每个干预都被表示为一个参与者级程序,描述目标参与者、预期行为、合法路线、交互顺序和时间约束,从而在异构基础模型之间实现统一的评估接口。TrafficImag评估四个互补的有效性维度:初始状态正确性、路线和行为有效性、交互一致性以及非目标保留,并且仅当所有四个维度都满足时,端到端的反事实才被视为成功。在最新的基础模型中,最强的推理器达到80.4%的宏F1分数,完整的条件接口将最佳生成器的端到端成功率从23.3%提高到55.0%。Oracle研究进一步表明,条件视频执行是主要剩余瓶颈。TrafficImag提供了一个可复现的基准,用于评估和诊断超越感知视频质量的反事实交通视频生成。

英文摘要

Existing roadside traffic datasets support perception, forecasting, and visual question answering, but they do not evaluate counterfactual video generation, in which a selected actor is modified and the generated future should remain consistent with road topology and unrelated traffic. We introduce TrafficImag, the first benchmark for counterfactual roadside traffic video generation. TrafficImag combines a large-scale roadside dataset (9,022 annotated images, 7,043 deduplicated video clips, and 31,145 actor-centered history-future samples) with an executable protocol that supports behavior reasoning, intervention-aware image editing, and conditional video generation. Each intervention is represented as an actor-level program describing the target actor, intended behavior, legal route, interaction order, and temporal constraints, enabling a unified evaluation interface across heterogeneous foundation models. TrafficImag evaluates four complementary validity dimensions: initial-state correctness, route and behavior validity, interaction consistency, and non-target preservation, and considers an end-to-end counterfactual successful only when all four are satisfied. Across state-of-the-art foundation models, the strongest reasoner reaches 80.4% macro F1, the complete condition interface raises end-to-end success from 23.3% to 55.0% for the best generator. Oracle studies further show that conditional video execution is the primary remaining bottleneck. TrafficImag provides a reproducible benchmark for evaluating and diagnosing counterfactual traffic video generation beyond perceptual video quality.

发表机构

  • The University of Texas at Austin(德克萨斯大学奥斯汀分校)
  • Duke University(杜克大学)

机构由 AI 辅助整理,请以论文原文为准。

↑