arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DataMagic:通过声明式多智能体编排创作数据视频

DataMagic: Authoring Data Videos through Declarative Multi-Agent Orchestration

Yupeng Xie, Zhenyang Wang, Liangwei Wang, Jiayi Zhu, Zhouan Shen, Yuyu Luo

arXiv 2609.33403首次发表:更新:

发表机构

The Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

DataMagic通过声明式多智能体编排,从原始表格数据自动生成数据视频,统一图表、解说与动画,提升质量与效率,优于现有LLM方法。

AI 中文摘要

数据视频通过动态图表、语音解说和同步动画来传达数据洞察,已成为一种广泛采用的数据叙事形式。然而,制作数据视频需要数据分析、叙事设计和视频编辑方面的专业知识。静态可视化工具缺乏叙事和动画能力;创作工具依赖于预先准备好的图表而非原始数据;像素级模型端到端生成视频,但无法保证数据准确性或来源可追溯性。端到端自动生成面临两个核心挑战:如何统一表示图表、解说和动画及其时间关系,以及如何高效搜索广阔的设计空间以找到叙事连贯的组成方案。我们提出了DataMagic,它通过声明式多智能体编排从原始表格数据创作数据视频。首先,声明式规范DVSpec将图表、解说和动画与数据绑定引用和声明式同步统一起来,确保数据来源可追溯和自动的视听对齐。其次,一种“先生成后编排”的多智能体策略并行生成候选场景,然后通过全局编排优化叙事连贯性。DVSpec为三种互补的交互模式提供了共享状态,弥合了全自动与细粒度人工控制之间的差距。对109个真实世界样本的评估表明,即使是最先进的LLM(例如GPT-5)也仅达到2.13/5的评分,执行成功率在48.62%到86.24%之间;DataMagic将质量提升至3.89(+83%),成功率超过95%,在动画和叙事维度上提升最为显著。一项用户研究表明,与对话式LLM工作流相比,DataMagic提高了创作效率(任务时间减少79.7%)并降低了感知认知负荷。项目页面:此https URL。

英文摘要

Data videos communicate data insights through dynamic charts, voice narration, and synchronized animations, and have become a widely adopted form of data storytelling. However, producing them requires expertise in data analysis, narrative design, and video editing. Static visualization tools lack narrative and animation capabilities; authoring tools rely on pre-prepared charts rather than raw data; and pixel-level models generate videos end-to-end but cannot guarantee data accuracy or provenance. End-to-end automatic generation faces two core challenges: how to uniformly represent charts, narration, and animations together with their temporal relationships, and how to efficiently search a vast design space for narrative-coherent compositions. We present DataMagic, which authors data videos from raw tabular data through declarative multi-agent orchestration. First, the declarative specification DVSpec unifies charts, narration, and animations with data-bound references and declarative synchronization, ensuring data provenance and automatic audio-visual alignment. Second, a "Generate-then-Orchestrate" multi-agent strategy generates candidate scenes in parallel and then optimizes narrative coherence through global orchestration. DVSpec provides a shared state for three complementary interaction modes, bridging full automation with fine-grained human control. Evaluations on 109 real-world samples show that even the most advanced LLM (e.g., GPT-5) achieves only 2.13/5 with execution success rates between 48.62% and 86.24%; DataMagic improves quality to 3.89 (+83%) with success rates above 95%, with the most significant gains in animation and narrative dimensions. A user study shows that, compared to a conversational LLM workflow, DataMagic improves creation efficiency (79.7% reduction in task time) and reduces perceived cognitive load. Project page: https://github.com/HKUSTDial/DataMagic.

CommentsAccepted at IEEE VIS 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑