arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PhysCaP:基于物理引导探索的代码即策略智能体

PhysCaP: Grounding Code-as-Policy Agent with Physics-Informed Exploration

Chen-Yu Lin, Jing-Wen Chen, Hsueh-En Chang, Hung-An Chen, Sheng-Hsun Chang, Chi-Pin Huang, Fu-En Yang, Min-Hung Chen, Yi-Ting Chen, Yu-Chiang Frank Wang, Shao-Hua Sun

arXiv 2608.21031首次发表:更新:

发表机构

National Taiwan University; NVIDIA Research; National Yang Ming Chiao Tung University(台湾大学; 英伟达研究院; 国立阳明交通大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PhysCaP是一种带物理引导探索层的代码即策略智能体,采用双智能体设计,可高效估算物体物理属性,在桌面操纵及LIBERO模拟任务中性能优于基线方法。

AI 中文摘要

我们提出了PhysCaP,一种用于机器人操纵主动感知的物理引导代码即策略(Code-as-Policy)智能体。尽管视觉-语言-动作策略在模仿演示方面表现出色,但它们依赖被动观察,无法推断操纵所需的潜在物理属性。PhysCaP在代码即策略框架中增加了物理引导探索层,可通过交互实现显式信息获取。它引入了无需训练的物理属性提取模块,该模块仅通过机器人本体感知即可估算物体质量和刚度,无需额外传感器。为平衡探索成本与获取信息的效率,PhysCaP采用双智能体设计:规划器(Planner)决定何时探索、何时停止;优先级器(Prioritizer)过滤不合理交互,并使用启发式优先级分数对剩余交互排序,从而实现高效、定向的探索。我们在真实桌面操纵任务(搜索隐藏物体、检测空罐子、寻找成熟牛油果)及LIBERO模拟任务中对PhysCaP进行评估。结果显示,现有被动及朴素交互基线要么在物理属性隐藏时失效,要么过度探索,而PhysCaP以更少的交互次数和更短的执行时间取得了相当的性能。消融研究进一步验证了所提物理属性提取模块的有效性。项目页面:this https URL

英文摘要

We present PhysCaP, a Physics-Informed Code-as-Policy agent system for active perception in robotic manipulation. While vision-language-action policies excel at imitating demonstrations, they rely on passive observation and fail to infer latent physical properties critical for manipulation. PhysCaP augments code-as-policy frameworks with a physics-informed exploration layer that enables explicit information-seeking through interaction. Our method introduces training-free physical property extraction modules that estimate object mass and stiffness from robot proprioception without additional sensors. To balance exploration costs and the efficiency of information obtained, PhysCaP employs a multi-agent design: a Planner that decides when to explore and when to stop, and a Prioritizer that filters implausible interactions and ranks the remainder using a heuristic priority score, enabling efficient, targeted exploration. We evaluate PhysCaP on three real-world tabletop manipulation tasks and a simulated task in LIBERO. The results show that existing passive and naive interactive baselines either fail when physical properties are hidden or over-explore, whereas PhysCaP achieves comparable performance with fewer interactions and reduced execution time. Ablation studies further validate the effectiveness of the proposed physical property extraction modules. Project page: https://physcap.github.io

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑