arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

AssemState:零样本家具装配的手册与物理状态引导推理

AssemState: Manual and Physical-State-Guided Reasoning for Zero-shot Furniture Assembly

Zhiyuan Qi, Jierui Li, Yifan Shen, Cheng Qian, Jiateng Liu

arXiv 2610.08446首次发表:更新:

发表机构

Tsinghua University; University of Illinois Urbana-Champaign; Xidian University(清华大学; 伊利诺伊大学厄巴纳-香槟分校; 西安电子科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对家具装配中3D空间推理难题,提出AssemState零样本框架,结合手册解析与物理状态反馈,显著提升装配树恢复和位姿精度。

AI 中文摘要

多模态大语言模型(MLLMs)在视觉理解方面取得了显著进展,但结合物理环境的精确3D空间推理仍然困难。家具装配不仅需要从图解手册中恢复步骤级操作,还需要将语义附着关系转化为6D位姿更新,使部件能够与环境和先前装配的组件进行物理交互。为研究此问题,我们提出AssemState,一个用于手册与物理状态引导的零样本家具装配框架。它首先采用锚点引导的边界装配状态将手册页面分解为单部件操作并恢复装配树。然后,它使用迭代的后状态反馈细化来指导连续的SE(3)更新和修正,并通过基于仿真的释放测试验证其物理合理性。实验表明,与最强先验基线相比,AssemState在装配树恢复上将F1从38.58%提升至62.80%,树精确匹配从28.24%提升至53.92%。在243个独立评估的部件级操作中,我们提出的迭代细化将评审接受的操作从0提升至5.3%,并将平均倒角距离从5.4111降至1.7744。然而,视觉上合理的候选位姿仍可能遭受碰撞、漂浮、镜像方向错误、不完全就位和错误侧附着等问题。这些结果表明,AssemState改善了操作结构恢复和选定的局部位姿指标,而MLLMs在空间关系推理方面仍存在局限。

英文摘要

Multimodal large language models (MLLMs) have made significant progress in visual understanding, but precise 3D spatial reasoning integrated with physical environment remains difficult. Furniture assembly requires not only recovering step-level operations from diagrammatic manuals, but also translating semantic attachment relations into 6D pose updates that enable parts to physically interact with the environment and previously assembled components. To study this problem, we propose AssemState, a zero-shot framework for manual and physical-state-guided furniture assembly. It firstly employs anchor-guided boundary assembly states to decompose manual pages into single-part operations and recover an assembly-tree. Then, it uses iterative after-state feedback refinement to guide successive (SE(3)) updates and corrections, and validates their physical plausibility through simulation-based release tests. Experiments show that compared with the strongest prior baseline, AssemState improves F1 from 38.58\% to 62.80\% and Tree Exact Match from 28.24\% to 53.92\% for assembly-tree recovery. On 243 independently evaluated part-level operations, our proposed iterative refinement improves judge-accepted operations from 0 to 5.3\% and reduces mean Chamfer distance from 5.4111 to 1.7744. However, visually plausible candidate poses may still suffer from collision, floating, mirror-orientation errors, incomplete seating, and wrong-side attachment. These results show that AssemState improves operation-structure recovery and selected local pose metrics, while MLLMs remain limited for spatial relationship reasoning.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑