arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2610.09781cs.ROcs.GR

端到端自主生成人类装配计划

End-to-End Autonomous Generation of Human Assembly Plans

Faustin Arion von Arx, Millicent Schlafly, Mark D. Fuge

首次发表
浏览论文内容

中文总结 AI 辅助

本研究提出一种端到端方法,将DfA原则编码进装配计划生成,仅需网格输入即可输出手册或失败报告,通过物理模拟和LLM实现工具清单、顺序、手册和反馈的自主生成,相比基线减少35%装配时间。

中文摘要 AI 辅助

将CAD设计转化为装配计划在很大程度上仍依赖人工完成,需要工程师考虑几何可行性、工具可达性、稳定性以及人类装配的人机工程学。在本工作中,我们将长期确立的面向装配的设计(DfA)原则编码为一种紧凑的端到端方法,用于生成装配计划。我们的方法仅需输入网格装配体,即可生成逐步装配手册或结构化失败报告,无需关节元数据、紧固件标注或额外信息。制造计划的四个主要组成部分被自主处理:装配工具清单、装配顺序与子装配体、装配手册,以及用于改进可装配性的设计反馈。为确定顺序计划,我们在物理模拟器中系统地拆卸物体,并应用编码了DfA原则的成本函数。手册生成、工具标注和装配反馈主要依赖多模态大语言模型。与Tian等人提出的始终先移除最外层部件的基线方法相比,DfA感知的顺序规划在136个包含5至30个部件的装配体上,以机器人臂装配时间代理测量的模拟装配时间减少了35%。在88.6%的装配步骤中选择了正确的工具。一个视觉-语言模型评判员将生成的手册与消融变体进行比较,识别哪些页面元素携带了读者所需的信息。所提出的方法和开源代码可供工程师或AI代理使用,以快速加速针对给定产品设计的制造计划创建。

英文摘要

Turning a CAD design into an assembly plan is still largely done by hand, requiring engineers to reason about geometric feasibility, tool access, stability, and the ergonomics of human assembly. In this work, we encode long-established design for assembly (DfA) principles into a contained, end-to-end approach for generating assembly plans. Our approach takes only a mesh assembly and produces either a step-by-step assembly manual or a structured failure report, requiring no joint metadata, fastener annotations, or additional information. Four major components of a manufacturing plan are addressed autonomously: an assembly tool list, the assembly sequence and subassemblies, an assembly manual, and design feedback for improving assemblability. For determining the sequence plan, we systematically disassemble the object in a physics simulator and apply a cost function that encodes DfA principles. Manual generation, tool labelling, and assembly feedback rely primarily on multimodal large language models. Compared with a baseline that always removes the outermost part first from Tian et al., DfA-aware sequence planning reduces simulated assembly time, measured with a robot-arm assembly-time proxy, by 35% on 136 assemblies of 5 to 30 parts. The correct tool is selected for 88.6% of assembly steps. A vision-language model judge compares the generated manuals against ablated variants, identifying which page elements carry the information a reader needs. The presented approach and open-source code are available for use by engineers or AI agents looking to rapidly accelerate the creation of manufacturing plans for a given product design.

发表机构

  • ETH Zurich(苏黎世联邦理工学院)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑