arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.16057cs.LGcs.AI

OmniHarness:通过符号策略学习实现通用视觉生成

OmniHarness: Harnessing Generalizable Visual Generation via Symbolic Policy Learning

  • Beihang University(北京航空航天大学)
  • The Chinese University of Hong Kong(香港中文大学)
  • National University of Singapore(新加坡国立大学)

机构由 AI 辅助整理,请以论文原文为准。

Xu Xu, Jinxiu Liu, Zhangbo Qiao, Jiaxing Lu, Xiangyu Zhang, Yubin Gu, Fangwei Ning, Yan Shi

AI总结:

OmniHarness通过符号策略学习实现通用视觉生成,将验证执行抽象为可复用策略,支持中间验证与自我导向练习,在六个基准上表现优异,显著提升解决率。

AI中文摘要:

统一多模态大语言模型(MLLMs)和多智能体系统推动了视觉生成的发展。然而,仍存在三个局限性:(1)现有方法通常提炼任务特定经验,泛化能力有限。(2)反思往往推迟到任务完成之后。(3)知识通常仅针对下游任务需求才被获取。为解决这些局限性,我们提出了OmniHarness,一个通过符号策略学习实现通用视觉生成的框架。OmniHarness将经过验证的执行过程抽象为面向视觉生成任务族的符号策略,捕获共享流程和适用条件,同时去除实例特定的输入。该框架(harness)为新的任务实例化、调整并组合这些策略。中间验证在执行过程中指导改进和失败恢复。通过自我导向的探究,OmniHarness在下游目标明确之前,自主生成并执行接近其能力极限的练习任务。执行反馈持续改进策略,而模型参数保持固定。在六个基准、三个MLLM骨干网络和三个视觉智能体框架上的实验表明,OmniHarness表现出强大的性能和持续的能力扩展。在ComfyBench的创意任务上,OmniHarness实现了95.0%的解决率,超出最强基线27.5个百分点。冻结的策略快照通过即插即用的复用改进了现有的视觉智能体系统。

英文摘要:

Unified multimodal large language models (MLLMs) and multi-agent systems have advanced visual generation. However, three limitations remain. (1) Existing methods often distill task-specific experience with limited generalizability. (2) Reflection is often deferred until task completion. (3) Knowledge is often acquired only in response to downstream task demands. To address these limitations, we introduce OmniHarness, a framework for generalizable visual generation via symbolic policy learning. OmniHarness abstracts verified executions into symbolic policies for visual generation task families, capturing shared procedures and applicability conditions while removing instance-specific inputs. The harness instantiates, adapts, and composes these policies for new tasks. Intermediate verification guides refinement and failure recovery during execution. Through self-directed inquiry, OmniHarness autonomously generates and executes practice tasks near its capability limits before downstream objectives are specified. Execution feedback continually refines the policies while model parameters remain fixed. Experiments across six benchmarks, three MLLM backbones, and three visual agent frameworks demonstrate strong performance and continual capability expansion. On ComfyBench's Creative tasks, OmniHarness achieves a 95.0% resolve rate, exceeding the strongest baseline by 27.5 percentage points. Frozen policy snapshots improve existing visual agent systems through plug-and-play reuse.

↑