presto:基于想象搜索的高效无训练开放世界物体放置方法
presto: Efficient, Training-free, and Open-world Object Placement via Imaginary Search
- Wuhan University(武汉大学)
- University of Louisville(路易斯维尔大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究提出零样本无训练框架presto,将开放世界物体放置转化为MLLM引导的启发式搜索任务,经多基准实验及人类研究验证,其在开放世界场景中性能最优,且MLLM作为评判者的变体更具感知一致性。
AI中文摘要:
物体放置是图像合成的关键,要求在多样场景中实现物体空间与语义一致的定位。现有方法通常依赖手工规则或在有限数据集上的监督学习,这限制了其泛化性和可解释性,尤其在涉及新物体和新场景的开放世界场景中。本研究将开放世界物体放置重新表述为多模态大语言模型(MLLM)推理引导的启发式搜索任务。我们提出了presto,这是一个零样本、无训练的框架,在想象动作空间内迭代优化物体位置和尺度。我们采用从粗到细的搜索策略确保快速收敛,并评估了两种决策变体:度量引导选择和MLLM作为评判者。在多个基准测试中的实验表明,presto达到了最先进的性能,特别是在之前未见过的开放世界设置中。人类研究进一步显示,MLLM作为评判者的变体比度量驱动方法产生更具感知一致性的放置,凸显了标准评估指标与人类视觉判断之间的差距。
英文摘要:
Object placement is critical in image composition, requiring spatially and semantically coherent positioning of objects within diverse scenes. Existing approaches typically rely on hand-crafted rules or supervised learning on limited datasets, which restricts their generalization and interpretability, especially in open-world scenarios involving novel objects and scenes. In this work, we reformulate open-world object placement as a heuristic search task guided by reasoning from a Multimodal Large Language Model (MLLM). We introduce \textsf{presto}, a zero-shot, training-free framework that operates within an imaginary action space to iteratively refine object position and scale. Our coarse-to-fine search strategy ensures fast convergence, and we evaluate two decision-making variants: Metric-guided Selection and MLLM-as-a-judge. Experiments across multiple benchmarks show that \textsf{presto}~achieves state-of-the-art performance, particularly in previously unseen, open-world settings. Human studies further reveal that the MLLM-as-a-judge variant produces more perceptually coherent placements than metric-driven approaches, highlighting a gap between standard evaluation metrics and human visual judgment.