arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29401cs.CV

OSEF:面向跨视频场景过程规划的一步证据融合

OSEF: One-Step Evidence Fusion for Cross-Video Scene Procedure Planning

Zhentong Ye, Lei Zhang, Sijia Zhou, Yingda Yu, Yuehan Shi, Jiaqi Xuan, Shuaiwu Dong, Guanchao Tong, Meimei Zhang, Bin Li

首次发表
浏览论文内容

中文总结 AI 辅助

针对跨视频场景过程规划的耦合障碍,构建含11个来源的基准,提出OSEF方法,在多个单元上提升规划成功率,表现优于现有SOTA。

中文摘要 AI 辅助

视频场景过程规划(VSPP)预先提供目标起始-目标观测结果,但规划器在证据需自行检索时应如何行动仍不明确。我们引入跨视频场景过程规划(CVSPP):给定答案被删改的起始-目标查询及K个候选视频,模型需检索支撑视频、定位相关窗口并预测动作序列。此处存在两个耦合障碍:相同任务演示共享阶段与窗口,且早期硬选择会向规划器传递错误场景链。我们构建了含11个来源、带类型负角色、采用闭式失败答案泄漏门及独立证据轴与规划轴指标的基准。在其14个源-视野单元上,我们适配9类规划器家族并以多数序列为基准。随后我们提出一步证据融合(OSEF),该方法对所有候选生成查询条件化的单元与跨度格,通过令牌全局适配器将完整软格输入规划器,且不预先裁剪任何窗口。OSEF在基准认定的6个可方法排名单元中均排名第一;在4个匹配的相同任务COIN与CrossTask单元上,其精确视频与规划成功率较增强型硬选择SOTA提升2.9至10.7个百分点,组件研究显示令牌全局接口贡献最大单增量;5个转换源单元处于或接近多数序列基准的基准上限。补充包包含模型构造器与评估代码。

英文摘要

Video Scene Procedure Planning (VSPP) supplies the target start-goal observations in advance, leaving open how a planner should act when the evidence must itself be retrieved. We introduce Cross-Video Scene Procedure Planning (CVSPP): given an answer-redacted start-goal query and K candidate videos, a model must retrieve the supporting video, localize the relevant window, and predict the action sequence. Two obstacles couple here. Same-task demonstrations share stages and windows, and an early hard selection passes the wrong scene chain to the planner. We build an eleven-source benchmark with typed negative roles, a fail-closed answer-leakage gate, and separate Evidence- and Plan-axis metrics. On its 14 source-horizon cells we adapt nine planner families against a majority-sequence floor. We then present One-Step Evidence Fusion (OSEF), which scores a query-conditioned cell-and-span lattice over all candidates and feeds the full soft lattice to the planner through a token-global adapter, cropping no window beforehand. OSEF ranks first on all six cells the benchmark certifies as method-rankable. On four matched same-task COIN and CrossTask cells it improves exact-video-and-plan success by 2.9-10.7 points over an enhanced hard-selection SOTA, and a component study assigns the largest single increment to the token-global interface. Five converted-source cells sit at or near the majority-sequence floor, the benchmark's remaining headroom. The supplementary package includes model constructors and evaluation code.

补充信息

↑