arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PACE-Bench:动态环境中基于代码演化的物理适应基准测试

PACE-Bench: Benchmarking Physics Adaptation via Code Evolution in Dynamic Environments

Yuhao Zhan, Bingxiang He, Zecong Tang, Chaojun Xiao

arXiv 2608.14441首次发表:更新:

发表机构

Tsinghua University; Zhejiang University(清华大学; 浙江大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PACE-Bench是涵盖六个物理领域144个适应对的模拟器支撑基准,用于测试智能体在环境物理条件变化时的代码演化适应能力,研究发现机制重新设计是核心瓶颈,现有方法性能仍有较大提升空间。

AI 中文摘要

自演化智能体可通过交互经验改进后续行为,但现有评估通常在固定执行条件下优化,未测试条件改变后的恢复能力。为解决这一差距,我们推出PACE-Bench(基于代码演化的物理适应),这是一个模拟器支撑的基准测试,涵盖六个物理领域的144个源到目标适应对。每对关联一个源环境和一个变异后的目标环境,二者目标与接口相同。在源环境中成功的代码驱动设计在目标环境中失效,智能体必须在有限尝试预算内,利用诊断沙盒反馈迭代将其调整为可在目标环境中运行的设计。我们比较了来自四个范式的十种自演化方法。该基准测试远未饱和:Reflexion + Qwen3-14B在完整基准测试对中仅成功35.9%,而GPT-5.5在完整预算下解决了静力学子集的66.7%。这些结果共同表明,模拟器支撑的反思比未经验证的自修正更可靠,记忆锚点将智能体绑定到早期设计,广泛的树搜索在不收敛的情况下进行探索。即使揭示精确的物理变化也未提升性能上限,这表明核心瓶颈是机制重新设计而非参数推断。数据和代码可在this https URL获取。

英文摘要

Self-evolving agents improve future behavior from interaction experience, yet existing evaluations typically optimize under fixed execution conditions and do not test recovery after those conditions change. To address this gap, we introduce PACE-Bench (Physics Adaptation via Code Evolution), a simulator-grounded benchmark of 144 source-to-target adaptation pairs across six physics domains. Each pair links a source environment to a mutated target environment with the same goal and interface. A code-driven design that succeeds in the source fails in the target, where agents must iteratively adapt it into a working target design using diagnostic sandbox feedback within a limited attempt budget. We compare ten self-evolving methods from four paradigms. The benchmark remains far from saturated: Reflexion + Qwen3-14B succeeds on only 35.9\% of full-benchmark pairs, while GPT-5.5 solves 66.7\% of the Statics subset under the full budget. Together, these results show that simulator-grounded reflection is more reliable than unverified self-revision, while memory anchors agents to early designs and broad tree search explores without converging. Even revealing exact physical changes does not raise the performance ceiling, pointing to mechanism redesign rather than parameter inference as the central bottleneck. Data and code are available at https://github.com/thunlp/PACE-Bench.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑