arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21543cs.CVcs.AI

presto:基于想象搜索的高效无训练开放世界物体放置方法

presto: Efficient, Training-free, and Open-world Object Placement via Imaginary Search

  • Wuhan University(武汉大学)
  • University of Louisville(路易斯维尔大学)

机构由 AI 辅助整理,请以论文原文为准。

Weixuan Ding, Shang Liu, Hanyu Pei, Zeyan Liu

AI总结:

本研究提出零样本无训练框架presto,将开放世界物体放置转化为MLLM引导的启发式搜索任务,经多基准实验及人类研究验证,其在开放世界场景中性能最优,且MLLM作为评判者的变体更具感知一致性。

AI中文摘要:

物体放置是图像合成的关键,要求在多样场景中实现物体空间与语义一致的定位。现有方法通常依赖手工规则或在有限数据集上的监督学习,这限制了其泛化性和可解释性,尤其在涉及新物体和新场景的开放世界场景中。本研究将开放世界物体放置重新表述为多模态大语言模型(MLLM)推理引导的启发式搜索任务。我们提出了presto,这是一个零样本、无训练的框架,在想象动作空间内迭代优化物体位置和尺度。我们采用从粗到细的搜索策略确保快速收敛,并评估了两种决策变体:度量引导选择和MLLM作为评判者。在多个基准测试中的实验表明,presto达到了最先进的性能,特别是在之前未见过的开放世界设置中。人类研究进一步显示,MLLM作为评判者的变体比度量驱动方法产生更具感知一致性的放置,凸显了标准评估指标与人类视觉判断之间的差距。

英文摘要:

Object placement is critical in image composition, requiring spatially and semantically coherent positioning of objects within diverse scenes. Existing approaches typically rely on hand-crafted rules or supervised learning on limited datasets, which restricts their generalization and interpretability, especially in open-world scenarios involving novel objects and scenes. In this work, we reformulate open-world object placement as a heuristic search task guided by reasoning from a Multimodal Large Language Model (MLLM). We introduce \textsf{presto}, a zero-shot, training-free framework that operates within an imaginary action space to iteratively refine object position and scale. Our coarse-to-fine search strategy ensures fast convergence, and we evaluate two decision-making variants: Metric-guided Selection and MLLM-as-a-judge. Experiments across multiple benchmarks show that \textsf{presto}~achieves state-of-the-art performance, particularly in previously unseen, open-world settings. Human studies further reveal that the MLLM-as-a-judge variant produces more perceptually coherent placements than metric-driven approaches, highlighting a gap between standard evaluation metrics and human visual judgment.

补充信息

↑