arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

更少令牌,更优行动:GPT-6 Astra 机器人智能体成功率提升14%且令牌消耗减少65%

Fewer Tokens, Better Action: GPT-6 Astra Robot Agents with 14% Higher Success Rate but 65% Fewer Tokens

Ruiyang Si, Jianxin Bi, Shunyu Yang, Rui Ni, Wenbo Huang, Qiang Wang, Shulong Jiang, Duomin Wang, Xiuyu Li, Haiwen Feng, Zhen Dong, Daquan Zhou

arXiv 2610.01939首次发表:更新:

发表机构

Peking University; National University of Singapore; NVIDIA; Impossible Research(北京大学; 新加坡国立大学; 英伟达; Impossible Research)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对VLM机器人智能体令牌开销大的问题,提出PyRUA-Lean框架,通过选择性观测和原语组合,在相等调用预算下将成功率从63.1%提至71.7%,并减少65%令牌。

AI 中文摘要

视觉语言模型(VLM)智能体可通过视觉反馈和动作原语控制机器人,但重复的模型调用和冗余的观测会产生大量令牌开销。我们提出PyRUA-Lean,一个交互式代码执行框架,将反馈驱动的原语组合与选择性观测相结合:智能体将经典机器人原语和学习到的视觉-语言-动作(VLA)策略组合成Python单元,这些单元执行条件检查与局部重试,仅返回明确请求的图像和用于重新规划的状态反馈。在来自LIBERO-PRO、RoboTwin 2.0和RoboCasa365的700个模拟任务实例中,我们将PyRUA-Lean与使用相同GPT-6 Astra规划器和底层机器人原语的工具调用基线进行比较。在相等的LLM调用预算下,PyRUA-Lean将整体成功率从63.1%提升至71.7%。在两个智能体均解决的实例中,它使用的LLM调用次数减少49%,输入令牌减少65%。

英文摘要

Vision language model (VLM) agents can control robots through visual feedback and action primitives, but repeated model invocations and redundant observations incur substantial token overhead. We introduce PyRUA-Lean, an interactive code-execution framework that couples feedback-driven primitive composition with selective observation: the agent composes classical robot primitives and learned vision-language-action (VLA) policies into Python cells that perform conditional checks and local retries, returning only explicitly requested images and state feedback for replanning. Across 700 simulated task instances from LIBERO-PRO, RoboTwin 2.0, and RoboCasa365, we compare PyRUA-Lean with a tool-calling baseline using the same GPT-6 Astra planner and underlying robot primitives. Under equal LLM-call budgets, PyRUA-Lean increases overall success from 63.1% to 71.7%. On instances solved by both agents, it uses 49% fewer LLM calls and 65% fewer input tokens.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑