arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.36593cs.GRcs.AI

Text2Sim:基于蒸馏专业知识的智能体物理仿真生成

Text2Sim: Agentic Physics-Based Simulation Generation with Distilled Expertise

Xiaoyu Xiong, Tsun-Hsuan Wang, Yi-Ling Qiao, Tao Du, Minchen Li

首次发表
浏览论文内容

中文总结 AI 辅助

Text2Sim提出一种基于Genesis的智能体流水线,利用分层结构和调试卡将文本请求自动转换为可执行物理仿真,在物理和视觉质量上超越四个最先进基线,并支持数据集构建等多模态扩展。

中文摘要 AI 辅助

创建多样化的物理仿真仍然是一项劳动密集型工作,因为资产、布局、物理参数、运动、控制和渲染必须联合设计和调试。我们提出Text2Sim,一个仿真专用的智能体流水线,可将纯文本请求转换为可执行、可编辑的动态案例。Text2Sim基于Genesis构建,采用分层智能体结构,结合了规划器与专门的编写器、资产生成工具和独立的评论家。从图形演示中蒸馏出的紧凑技能(调试卡)为基于执行的修复提供角色特定的物理指导。我们在42个保留提示上评估物理质量、视觉质量和人类偏好,涵盖刚体、铰接、可变形和布料现象,并在经验构建和评估之间进行论文级别的划分。我们设计了自动物理和视觉评分器来评估结果质量,Text2Sim在这两项指标上都取得了比所有四个最先进基线更高的分数。在与这些基线的盲法用户研究中,明显更多的参与者偏好Text2Sim而非基线,这与我们自动评分器的结果一致。该流水线还支持广泛的 downstream 应用;我们选择数据集构建和扩展到多模态输入作为两个代表性示例。我们将发布代码、调试卡库和生成案例数据集,每个案例将文本提示和渲染视频与可执行程序、资产、物理参数、控制和记录状态配对。

英文摘要

Creating diverse physical simulations remains labor-intensive because assets, layout, physical parameters, motion, control, and rendering must be designed and debugged jointly. We present Text2Sim, a simulation-specialized agentic pipeline that converts a text-only request into an executable, editable dynamic case. Built on Genesis, Text2Sim uses a hierarchical agentic structure that combines a Planner with specialized Writers, asset-generation tools, and an independent Critic. Compact skills (Debug Cards) distilled from graphics demonstrations provide role-specific physical guidance for execution-based repair. We evaluate physical quality, visual quality, and human preference on 42 held-out prompts spanning rigid, articulated, deformable, and cloth phenomena, with a paper-level split between experience construction and evaluation. We design automatic physical and visual scorers to evaluate the quality of the results, and Text2Sim achieves higher scores than all four state-of-the-art baselines on both metrics. In blinded user studies with these baselines, significantly more participants prefer Text2Sim than prefer the baselines, which is consistent with the results from our automatic scorers. The pipeline also supports a broad range of downstream applications; we select dataset construction and extension to multimodal input as two representative examples. We will release the code, the Debug Card library, and a dataset of generated cases, each pairing the text prompt and rendered video with the executable program, assets, physical parameters, controls, and recorded states.

↑