arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FoldingAgent:从演示视频推断参数化折纸流程的智能体框架

FoldingAgent: Inferring Parametric Origami Procedures from Demonstration Videos

Maya Moriya, Sigal Raab, Yael Vinker, Tali Dekel

arXiv 2609.00377首次发表:更新:

发表机构

Weizmann Institute of Science; MIT(魏茨曼科学研究所; 麻省理工学院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出FoldingAgent智能体框架,结合VLM推理与专用工具及物理模拟,可从折纸演示视频推断参数化折纸流程,缓解多步折纸的累积误差,缩小人类折纸知识与计算方法的差距。

AI 中文摘要

我们提出了FoldingAgent,这是一种智能体框架,可直接从折纸演示视频中推断出明确的参数化折纸程序。该框架利用预训练视觉语言模型(VLM)的推理能力,并配备了一套专用工具,使智能体能够模拟几何转换、验证物理合理性、检索和比较视觉内容,并评估自身的预测结果。为了将视觉内容转换为折纸程序,我们定义了一个参数空间,该空间包含纸张的几何形状和一组参数化折纸动作。与预测静态折痕图案的模型不同,我们的智能体按顺序运行,具备重新规划动作的能力,有效缓解了多步折纸中固有的累积误差问题。我们的研究旨在缩小人类折纸知识与计算方法之间的差距:人类折纸知识主要通过非结构化的视觉演示进行传递,而计算方法通常依赖结构化的参数化表示,如折痕图案或可执行的参数化计划。我们在PurelandFold上对所提方法进行了评估,这是一个新整理的基准,包含各种Pureland折纸视频,带有真实几何形状和动作标签。结果表明,通过将VLM推理与专用工具及物理模拟相结合,我们能够成功将非结构化视觉演示转换为可执行、物理上合理的折纸流程。

英文摘要

We present FoldingAgent, an agentic framework for inferring explicit parametric folding programs directly from origami demonstration videos. Our framework leverages the reasoning power of a pre-trained Vision-Language Model (VLM) equipped with a suite of specialized tools that enable the agent to simulate geometric transitions, verify physical plausibility, retrieve and compare visual content, and evaluate its own predictions. To translate visual content into folding programs, we define a parametric space that consists of the paper's geometry and a set of parametric folding actions. Unlike models that predict static crease patterns, our agent operates sequentially and possesses the ability to re-plan its actions, effectively mitigating the compounding errors inherent in multi-step folding. Our approach takes a step toward closing the gap between human origami knowledge, which is primarily shared through unstructured visual demonstrations, and computational methods, which typically rely on structured, parametric representations such as a crease pattern or an executable parametric plan. We evaluate our approach on PurelandFold, a newly curated benchmark of diverse Pureland origami videos with ground-truth geometry and action labels. Our results demonstrate that by combining VLM reasoning with a set of specialized tools and physical simulation, we can successfully transform unstructured visual demonstrations into executable, physically plausible folding procedures.

CommentsProject Page: https://maya-moriya.github.io/origami-page/ Accepted to SIGGRAPH ASIA 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

相关深度报道

↑