发表机构
Peking University; SenseTime Research; Zhejiang University(北京大学; 商汤科技研究院; 浙江大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有视觉上下文学习对物理基础图像变换支持不足的问题,提出TransPhy框架,通过分解为物理规则归纳和过渡对齐渲染,在PhysVICL-74基准上提升了相关性能。
AI 中文摘要
视觉演示为指定难以用文字详尽描述的图像变换提供了自然界面。然而,现有视觉上下文学习(VICL)方法主要关注外观层面的关系迁移,对物理基础变换的支持有限,这类变换的结果取决于材料属性、几何形状、物体交互和环境条件。给定源-目标示例对和查询图像,面向物理基础的VICL要求模型推断演示的变换、将其效果适配到查询特定的场景上下文,并保留与规则无关的内容。我们推出PhysVICL-74,包含74条物理基础变换规则和5240对源-目标图像,构成近75000个训练与评估上下文。其基准划分分别评估新实例迁移和未见过规则的泛化能力。我们进一步提出TransPhy框架,将面向物理基础的VICL分解为物理规则归纳和过渡对齐渲染两个阶段。TransPhy首先预测演示的规则和明确的查询特定目标状态描述,再通过逐令牌的专家混合适配合成目标图像,专家路由由局部过渡线索引导。实验表明,TransPhy在物理规则 adherence( adherence 保留英文)、查询一致性和未见过规则的泛化能力方面优于现有视觉上下文编辑方法。
英文摘要
Visual demonstrations provide a natural interface for specifying image transformations that are difficult to describe exhaustively with text. However, existing visual in-context learning (VICL) methods primarily focus on appearance-level relation transfer and provide limited support for physically grounded transformations, whose outcomes depend on material properties, geometry, object interactions, and environmental conditions. Given a source--target exemplar pair and a query image, physically grounded VICL requires a model to infer the demonstrated transformation, adapt its effects to the query-specific scene context, and preserve rule-irrelevant content. We introduce PhysVICL-74, comprising 74 physically grounded transformation rules and 5,240 source--target image pairs that form nearly 75K training and evaluation contexts. Its benchmark split separately evaluates novel-instance transfer and unseen-rule generalization. We further propose TransPhy, a framework that decomposes physically grounded VICL into physical-rule induction and transition-aligned rendering. TransPhy first predicts the demonstrated rule and an explicit query-specific target-state description, and then synthesizes the target image through token-wise mixture-of-experts adaptation, with expert routing guided by localized transition cues. Experiments show that TransPhy improves physical-rule adherence, query consistency, and unseen-rule generalization over existing visual in-context editing methods.