arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OmniPhys:文本到图像生成中物理常识的知识图谱驱动基准测试与集体优化

OmniPhys: Knowledge-Graph-Driven Benchmarking and Collective Optimization for Physical Commonsense in Text-to-Image Generation

Yajing Xu, Yarong Lan, Jiaoyan Chen, Yichi Zhang, Jeff Z. Pan, Mingchen Tu, Zhizhen Liu, Wen Zhang, Huajun Chen

arXiv 2607.25641首次发表:更新:

发表机构

Zhejiang University; The University of Manchester; The University of Edinburgh; Ant Group(浙江大学; 曼彻斯特大学; 爱丁堡大学; 蚂蚁集团)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对文本到图像模型常违反物理常识及现有基准测试不足等问题,引入基于物理知识图谱的OmniPhys基准测试,提出OmniPrompt迭代框架,经对12个模型评估,显著提升了物理一致性。

AI 中文摘要

虽然文本到图像模型具有很高的视觉保真度,但它们经常违反基本的物理常识。现有基准测试往往依赖于粗粒度描述,无法诊断对特定物理原理的掌握情况。此外,生成过程的高随机性导致当前的提示优化方法存在梯度幻觉问题。为应对这些挑战,我们引入了OmniPhys,一个基于物理知识图谱的包含1551个样本的严格基准测试。通过将PhET模拟与标准课程对齐,OmniPhys实现了一个知识到场景的管道,通过双路径验证协议进行诊断压力测试。我们还提出了OmniPrompt,一个将物理对齐视为离散优化问题的迭代框架。评估显示了通用的物理瓶颈,结果表明OmniPrompt显著提高了不同骨干模型的物理一致性。

英文摘要

While text-to-image models exhibit remarkable visual fidelity, they frequently violate fundamental physical commonsense. Existing benchmarks often rely on coarse-grained descriptions, failing to diagnose the mastery of specific physical principles. Moreover, the high stochasticity of generative processes causes current prompt optimization methods to suffer from gradient hallucinations, where optimizers are misled by transient visual artifacts rather than systemic flaws. To address these challenges, we introduce OmniPhys, a rigorous benchmark of 1,551 samples grounded in a Physical Knowledge Graph. By aligning PhET simulations with standard curricula, OmniPhys operationalizes a knowledge-to-scenario pipeline that performs diagnostic stress tests via a dual-path verification protocol. We further propose OmniPrompt, an iterative framework that treats physical alignment as a discrete optimization problem. For each query, OmniPrompt aggregates K stochastic images into a per-query feedback buffer. Across training, it further merges feedback from batches of B queries before each meta-policy update, filtering seed and query-local noise. Evaluations across 12 representative text-to-image models reveal universal physical bottlenecks. Results demonstrate that OmniPrompt significantly enhances physical consistency across diverse backbones, proving the transferability and efficacy of our evolved meta-policies. The code and data are available at https://github.com/zjukg/OmniPhys

Commentsaccepted by KDD 2026 DB track

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑