arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越可执行模型:用于物理系统建模的Pufibara智能体框架与Modelica智能体工作流基准

Beyond Executable Models: The Pufibara Agent Harness and the Modelica Agent Workflow Benchmark for Physical System Modeling

Zizhe Wang

arXiv 2608.23653首次发表:更新:

AI 中文总结

该研究针对物理系统建模中智能体的状态管理与证据关联问题,提出Pufibara智能体框架,并构建含232个任务的基准,实验显示其在任务成功率与资源效率上优于Claude Code。

AI 中文摘要

AI智能体越来越多地用于仿真驱动的工程领域。物理系统建模与软件工程中的通用代码生成有着不同的要求,因为其正确性不仅取决于语法和可执行性,还取决于物理一致性和依赖场景的行为。我们在Modelica(一种基于方程的建模语言)中研究这一挑战,在该语言中,模型可能能够编译和仿真,但仍可能违反其预期的物理或工程要求。在连续的修订过程中,智能体可能会迷失需求,或依赖过时候选模型产生的仿真证据。为解决这一挑战,我们提出Pufibara——一种智能体框架,它可在各修订版本间维护持久的工程状态,将执行和仿真证据与产生该证据的候选模型关联,并将提交设为显式的智能体动作。为评估端到端的Modelica智能体工作流,我们还提出一种基于源代码的方法,用于构建现实且可独立评估的任务。我们使用该方法构建了包含232个任务的Modelica智能体工作流基准,涵盖模型修复、模型生成和模型调优。每个提交的候选模型由基准拥有的评估器在智能体循环之外进行评分。我们在两个匹配的大语言模型(LLM)后端下,将Pufibara作为完整框架与Claude Code进行对比。使用DeepSeek v4 Flash时,Pufibara通过202个任务,而Claude Code通过185个;使用Claude Sonnet 5时,Pufibara通过202个任务,Claude Code通过187个。根据仓库报告的token统计,Pufibara的逻辑token总量降低了76.4%-82.5%,其顺序运行时间降低了6.1%-58.4%。这些结果表明,即使在匹配的LLM后端下,完整的智能体框架在物理系统建模的任务成功率和资源使用上也存在显著差异。

英文摘要

AI agents are increasingly used for simulation-driven engineering. Physical system modeling presents different requirements from general-purpose code generation in software engineering, because correctness depends not only on syntax and executability but also on physical consistency and scenario-dependent behavior. We study this challenge in Modelica, an equation-based modeling language in which a model may compile and simulate while still violating its intended physics or engineering requirements. Across successive revisions, an agent may lose track of requirements or rely on simulation evidence produced by an outdated candidate. To address this challenge, we present Pufibara, an agent harness that maintains persistent engineering state across revisions, associates execution and simulation evidence with the candidate that produced it, and makes submission an explicit agent action. To evaluate end-to-end Modelica agent workflows, we also propose a source-grounded method for constructing realistic and independently evaluable tasks. We use this method to build the 232-task Modelica Agent Workflow Benchmark, spanning Model Repair, Model Generation, and Model Tuning. Each submitted candidate is scored by a benchmark-owned evaluator outside the agent loop. We compare Pufibara with Claude Code as complete harnesses under two matched large language model (LLM) backends. With DeepSeek v4 Flash, Pufibara passes 202 tasks, compared with 185 for Claude Code. With Claude Sonnet 5, Pufibara passes 202 tasks, compared with 187 for Claude Code. Under the repository-reported token accounting, Pufibara records 76.4%-82.5% lower logical-token totals. Its sequential runtime is 6.1%-58.4% lower. These findings show that, even under matched LLM backends, complete agent harnesses can differ substantially in both task success and resource use for physical system modeling.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑