arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AVA-Encoder:面向智能体原生的视频表示学习

AVA-Encoder: Towards Agent-Native Video Representation Learning

Chuyue Li, Jinpeng Yu, Haozhe Wang, Tian Xueyun, Zhijing Zhang, Bingnan Li, Shuqi Gu, Kan Ren, Jiaming Liu, Ruihua Huang

arXiv 2608.12313首次发表:更新:

发表机构

Qwen Business Unit of Alibaba; ShanghaiTech University; The Hong Kong University of Science and Technology; Institute of Computing Technology; Southeast University(阿里巴巴通义千问业务部; 上海科技大学; 香港科技大学; 中国科学院计算技术研究所; 东南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对创意智能体缺乏从高质量电影学习的有效方式的问题,提出AVA-Encoder框架,其视频表示经重构优化后,在相关基准上性能优于现有方法,还发布了配套框架、基准及数据集。

AI 中文摘要

创意智能体仍缺乏从高质量人类电影中学习的有效方式,这限制了它们生成电影级视频的能力。一个关键挑战在于缺少既忠实于电影内容又能直接用于智能体推理与操作的结构化视频表示。为解决该挑战,我们提出了Agentic Video Auto-Encoder(AVA-Encoder,智能体式视频自动编码器),这是一种通过智能体自动编码学习智能体原生视频表示的框架。AVA-Encoder将视频转换为知识图谱(KG)表示,再重构回视频;其层级节点与状态节点存储结构化文本,链接资产层则保存生成的图像、音频和视频,类型化边以智能体易于理解、查询和编辑的形式保留这些文本描述与资产间的关系。视频重构差异驱动文本梯度优化框架,该框架将评估反馈表达为自然语言更新方向,用于外层循环中与数据无关的编码策略伪训练,以及测试时内层循环中可选的与数据相关的KG表示优化。大量实验表明,AVA-Encoder比最强的外部基准模型提升了20.7个百分点;在仅策略的受控设置中,其伪训练的镜头级Agentic Video Encoder策略还优于精心人工调优的策略,同时使用的系统提示词令牌减少了74.3%。我们发布了完整的AVA-Encoder框架、可靠的智能体式视频重构基准,以及首个高质量电影KG表示数据集。

英文摘要

Video creative agents still lack an effective way to learn from high-quality human films, limiting their ability to produce cinematic-grade videos. A key challenge is the absence of a structured video representation that is both faithful to film content and directly usable for agentic reasoning and manipulation. To address the challenge, we propose the Agentic Video Auto-Encoder (AVA-Encoder), a novel auto-encoding framework driven by agentic self-evolution to learn agent-native video representations. AVA-Encoder transforms a video into a Film Knowledge Graph (KG) representation and then reconstructs it back into video. This Film KG representation explicitly captures entities, events, assets, and their multimodal relationships in a structured form that can be easily understood, queried, and manipulated by agents. The reconstruction residual drives a dual-loop textual-gradient optimization framework that jointly improves the Film KG representation and the Agentic Video Encoder. Extensive experiments show that AVA-Encoder achieves a 20.7-percentage-point absolute gain, or a 73.1% relative improvement, over the strongest external baseline. In the controlled policy-only setting, its pseudo-trained Agentic Video Encoder policy also outperforms a carefully human-tuned policy while using 74.3% fewer shot-level and 70.1% fewer keyframe-level system-prompt tokens. We release the complete AVA-Encoder framework, a reliable agentic video reconstruction benchmark, and the first dataset of high-quality Film KG representations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑