arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MaLiang-Harness:通往图像与视频生成的可编程路径

MaLiang-Harness: A Programmable Path to Image and Video Generation

Haoyu Zhao, Zihao Zhang, Xudong Wang, Jiaxi Gu, Zuxuan Wu, Yu-Gang Jiang, Shuicheng Yan

arXiv 2609.34309首次发表:更新:

发表机构

National University of Singapore; Fudan University; Tencent(新加坡国立大学; 复旦大学; 腾讯)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对程序正确执行但视觉输出不符的问题,提出MaLiang-Harness框架,通过持久状态、可追踪过程和修订验证机制,实现MLLM驱动的图像与视频生成,实验表明GPT-6-Astra在生成成功率上表现优异。

AI 中文摘要

可执行程序为图像和视频的构建方式提供了显式控制,但生成可运行的代码仅仅是视觉创作的开始。程序可能正确执行,却违反了所要求的构图、外观或运动。我们将这种差异定义为程序到视觉(P2V)差距,并引入MaLiang-Harness,一个统一框架,将基于多模态大语言模型(MLLM)的视觉生成组织为构建、检查和修订的持续过程。其核心设计是使不断演化的视觉程序、其构建历史及其验证共享一个共同的修订参考。我们将持久可执行生成(PEG)状态定义为保留程序和任务上下文。可追踪生成过程(TGP)将编辑与渲染证据联系起来,而修订感知的编辑与验证(REV)支持恢复,并在完成前检查当前修订。这些机制共同协调跨渲染后端的规划、执行和视觉反馈。我们在MaLiang-IBench上评估了11个强大的闭源MLLM,在MaLiang-VBench上评估了4个,衡量生成成功率、视觉质量和计算成本。GPT-6-Astra在两个基准上均达到100%的生成成功率,其中96.0%的图像任务和76.9%的视频任务满足所有质量阈值。比较还揭示了通用能力分数与视觉生成性能之间的不匹配,得分相似的模型在满足视觉要求的能力上存在显著差异。MaLiang-Harness为研究MLLM如何将可执行代码转化为视觉结果提供了系统基础,既展示了可编程生成的潜力,也揭示了通用基准作为预测指标的局限性。项目可在该https URL获取。

英文摘要

Executable programs offer explicit control over how images and videos are constructed, but generating runnable code is only the beginning of visual creation. A program can execute correctly while violating the requested composition, appearance, or motion. We define this discrepancy as the Program-to-Visual (P2V) gap and introduce MaLiang-Harness, a unified framework for organizing MLLM-driven visual generation into a persistent process of construction, inspection, and revision. Its central design is to make the evolving visual program, its construction history, and its verification share a common revision reference. We define the Persistent Executable Generation (PEG) state as preserving programs and task context. Traceable Generation Process (TGP) connects edits to rendered evidence, and Revision-aware Editing and Verification (REV) supports restoration and checks the current revision before completion. Together, these mechanisms coordinate planning, execution, and visual feedback across rendering backends. We evaluate 11 powerful closed-source MLLMs on MaLiang-IBench and four on MaLiang-VBench, measuring generation success, visual quality, and computational cost. GPT-6-Astra achieves 100% generation success on both benchmarks, with 96.0% of image tasks and 76.9% of video tasks meeting all quality thresholds. The comparison also reveals a mismatch between general capability scores and visual generation performance, with similarly scored models differing substantially in their ability to satisfy visual requirements. MaLiang-Harness provides a systematic basis for studying how MLLMs translate executable code into visual outcomes, exposing both the potential of programmable generation and the limitations of general benchmarks as predictors of this ability. The project is available at https://github.com/gulucaptain/MaLiang-Harness.

Comments22 pages, 11 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑