AgenticCADedit:一种有状态、工具介导的多模态3D CAD编辑智能体方法
AgenticCADedit: A Stateful, Tool-Mediated Agentic Approach to Multimodal 3D CAD Editing
浏览论文内容
中文总结 AI 辅助
针对现有神经CAD编辑方法无状态、无法累积进度的问题,提出AgenticCADedit,将编辑分解为有状态CAD上的增量可验证操作,提升三个LLM的编辑有效性与接受率,并降低令牌成本。
中文摘要 AI 辅助
计算机辅助设计是工业制造的核心,设计师的日常工作很大一部分涉及根据包含语音、草图和模型交互的多模态请求来编辑现有模型。现有的神经CAD方法主要侧重于无条件或文本条件生成。neuralCAD-Edit方法形式化了专家多模态编辑请求,但其迭代基线在多次尝试中完善完整的CAD程序,并在无状态CAD环境中从原始模型执行每次尝试。因此,每次尝试都必须从头重建整个编辑,部分正确的进度被丢弃而非累积,模型既无法检查刚生成的几何体,也无法选择性地回退单个错误操作。我们提出AgenticCADedit,它将编辑转化为对持久CAD状态的一系列小而可验证的操作,而非单个重新生成的程序。它不是生成一个完整程序,而是应用增量代码步骤,每个步骤都提交到会话,检查生成的面和边,渲染高亮选择以验证目标区域已被处理,并在未处理时回退单个操作。因此,后续操作基于先前操作生成的几何体。我们的方法在三个评估的LLM(开放权重:qwen3.6-27b、gemma4-31b;专有:gpt-5.6-luna)的所有指标上均有提升,其中对最弱的基线模型qwen3.6-27b的提升最大,其有效性从51.0%提高到94.8%,接受率从1.6%提高到12.0%。使用gpt-5.6-luna进行的令牌成本分析进一步显示,输出令牌比neuralCAD-Edit减少66.7%,而94.8%的输入令牌由提示缓存提供。
英文摘要
Computer-aided design is central to industrial manufacturing, and much of a designer's daily work consists of editing existing models from multimodal requests involving speech, sketches, and model interaction. Existing neural CAD approaches focus predominantly on unconditional or text-conditioned generation. The neuralCAD-Edit approach formalizes expert multimodal editing requests, but its iterative baseline refines a complete CAD program across attempts, executing each attempt from the original model in a stateless CAD environment. Every attempt must therefore reconstruct the entire edit from scratch, so partially correct progress is discarded rather than accumulated, and the model can neither inspect the geometry it has just produced nor selectively revert a single faulty operation. We present AgenticCADedit, which turns editing into a sequence of small, verifiable actions on a persistent CAD state instead of a single regenerated program. Rather than emitting one complete program, it applies incremental code steps that each commit to the session, inspects the resulting faces and edges, renders highlighted selections to verify that the intended region was addressed, and reverts individual operations when it was not. Subsequent actions therefore build on the geometry produced by earlier ones. Our approach improves on all metrics for all three evaluated LLMs (open-weight: qwen3.6-27b, gemma4-31b; proprietary: gpt-5.6-luna), with the largest gains for the weakest baseline model, qwen3.6-27b, whose validity rises from 51.0% to 94.8% and acceptance from 1.6% to 12.0%. A token-cost analysis with gpt-5.6-luna further shows $66.7$% fewer output tokens than neuralCAD-Edit, while $94.8$% of input tokens are served from the prompt cache.
发表机构
- Fraunhofer IGD(弗劳恩霍夫计算机图形研究所)
- TU Darmstadt(达姆施塔特工业大学)
- Delft University of Technology(代尔夫特理工大学)
机构由 AI 辅助整理,请以论文原文为准。