arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

VIS-Ground:具有上下文锚定的视频交互式叙事

VIS-Ground: Video Interactive Storytelling with Contextual Grounding

Bingxuan Li, Yiwen Song, Xueqing Wu, Yanzhou Pan, Yang Li, Kuang Su, Jingyun Liu, Sebastian Ko, Huan Zhang, Tong Zhang, Nanyun Peng, Tomas Pfister, Yale Song

arXiv 2610.09326首次发表:更新:

发表机构

Google; University of Illinois at Urbana-Champaign; University of California, Los Angeles(谷歌; 伊利诺伊大学厄巴纳-香槟分校; 加利福尼亚大学洛杉矶分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

VIS-Ground提出将异构输入转化为可执行约束模型,通过结构化抽象、约束归纳和约束生成,在多个骨干网络上平均提升10.3分,解决交互式视频生成中的上下文锚定问题。

AI 中文摘要

视频交互式叙事使观众能够主动引导视频的展开方式。然而,一旦允许观众在生成过程中进行干预,就会出现一个新的挑战:观众的请求可能对锚定源和当前渲染的视频状态都存在潜在依赖。这些依赖可能不会在任何单个输入中明确表述,而只有在源、渲染历史和新的观众意图被联合考虑时才会显现。现有的交互式视频生成系统主要强调遵循观众指令,而基于源的视频生成方法则侧重于将生成内容与外部叙事或知识源对齐。这留下了一个基本问题未被充分探索:在交互式延续过程中,生成模型应锚定在什么上下文上,以及如何将异构、非结构化的输入转换为这样的锚定上下文?在这项工作中,我们将上下文锚定定义为将异构输入上下文转换为用于视频生成的可执行约束模型的过程。为了解决这一挑战,我们引入了VIS-Ground,它执行结构化上下文抽象以恢复锚定状态和跨上下文依赖,生成约束归纳以将相关依赖投射到特定于候选的约束中,以及约束视频生成以通过规划、验证、修订和渲染来强制执行这些约束。在三个视频生成骨干网络上,VIS-Ground始终达到最高的整体综合得分,相对于每个骨干网络的最强基线,平均绝对改进达到10.3分。详细分析进一步显示了在叙事和知识锚定方面的增益,并揭示了在依赖提取和视频渲染过程中的忠实实现方面仍然存在的挑战。

英文摘要

Video interactive storytelling enables viewers to actively steer how a video unfolds. However, once we allow viewers to intervene during generation, a new challenge arises: The viewer's request can have latent dependencies on both the grounding source and the current rendered video state. These dependencies may not be explicitly stated in any individual input, but emerge only when the source, rendered history, and new viewer intent are considered jointly. Existing interactive video generation systems primarily emphasize following viewer instructions, while source-grounded video generation methods focus on aligning generated content with an external narrative or knowledge source. This leaves a fundamental question underexplored: What context should a generation model ground on during interactive continuation, and how can heterogeneous, unstructured inputs be transformed into such grounding context? In this work, we formulate contextual grounding as the process of transforming heterogeneous input context into an executable constraint model for video generation. To address this challenge, we introduce VIS-Ground, which performs Structured Context Abstraction to recover grounded states and cross-context dependencies, Generation Constraints Induction to project relevant dependencies into candidate-specific constraints, and Constrained Video Generation to enforce these constraints through planning, verification, revision, and rendering. Across three video generation backbones, VIS-Ground consistently achieves the highest overall composite score, reaching an average absolute improvement of 10.3 points over the strongest per-backbone baselines. Detailed analysis further shows gains across both narrative and knowledge grounding, and reveals remaining challenges in dependency extraction, and faithful realization during video rendering.

CommentsProject Page: https://bx126.github.io/vis-ground.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑