arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ScaffoldM3C:一种用于生成式稳定施工规划的多模态序贯蒙特卡洛框架

ScaffoldM3C: A Multimodal Sequential Monte Carlo Framework for Generative Stable Construction Planning

Gadiel Sznaier Camps, Chengyang He, Guillaume Sartoretti, Eduardo Montijano, Mac Schwager

arXiv 2610.00487首次发表:更新:

发表机构

Stanford University; National University of Singapore; Universidad de Zaragoza(斯坦福大学; 新加坡国立大学; 萨拉戈萨大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ScaffoldM3C提出多模态序贯蒙特卡洛框架,通过脚手架标记和候选块采样实现稳定施工规划,模型小4倍、推理快5-20倍,质量相当且稳定性更高。

AI 中文摘要

自主构建物理上可实现的三维结构仍然是一个重大挑战,原因在于组合动作空间、可互换组件、等效装配序列以及施工过程中严格的稳定性要求。最先进的方法通过微调大型语言模型来实现基于文本的生成式施工。然而,这些方法不允许多模态(文本、图像、草图)条件设定,忽视了脚手架在稳定中间结构中的实际作用,并且推理速度较慢。因此,我们将施工问题表述为具有多种潜在装配动作和多种潜在任务条件模态的概率性下一块生成任务。同时,我们通过引入辅助脚手架块标记来显式考虑脚手架的效用。我们提出了Scaffold多模态蒙特卡洛(ScaffoldM3C)框架,这是一种用于基于块的稳定施工的多模态、轻量级自回归模型,它提出一组下一步候选块。利用这些候选块,我们采用序贯蒙特卡洛(SMC)方法来维护一组可能的装配序列,从而允许我们同时考虑多个可能不同的装配方向。我们通过扩展StableText2Brick数据集以包含图像条件提示和脚手架稳定的构建序列来训练我们的多模态架构。ScaffoldM3C比竞争基线模型小4倍,在推理过程中实现了5倍到20倍的加速,同时达到了与最先进方法相当的施工质量和更高的整体稳定性。我们通过模拟和真实机器人装配演示证明了我们方法的有效性。

英文摘要

Autonomously constructing physically realizable 3D structures remains a significant challenge due to combinatorial action spaces, interchangeable components, equifinal assembly sequences, and strict stability requirements during construction. State-of-the-art methods fine-tune large language models for text-based generative construction. However, these approaches do not allow for Multimodal (text, image, sketch) conditioning, overlook the practical role of scaffolding for stabilizing intermediate structures, and suffer from slow inference speeds. Therefore, we formulate construction as a probabilistic next-block generation task with multiple potential assembly actions and multiple potential task conditioning modalities. Concurrently, we explicitly consider the utility of scaffolding by introducing an auxiliary scaffold block token. We present Scaffold Multimodal Monte Carlo (ScaffoldM3C), a multimodal, lightweight, auto-regressive model for stable block-based construction, that proposes a set of next-step candidate blocks. Leveraging these candidates, we utilize Sequential Monte Carlo (SMC) to maintain a population of possible assembly sequences, allowing us to consider multiple, potentially different, assembly directions simultaneously. We train our multimodal architecture by extending the StableText2Brick dataset to contain image conditioning prompts and scaffold-stabilized build sequences. ScaffoldM3C is 4x smaller than competing baselines, yielding a 5x to 20x speedup during inference, while achieving comparable construction quality to state-of-the-art methods and higher overall stability. We demonstrate the effectiveness of our approach through simulations and real-world robot assembly demonstrations.

CommentsThis work has been submitted to IEEE Transactions on Robotics and Learning (T-RL) and is currently under review. Project Page: https://stanfordmsl.github.io/ScaffoldM3C/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑