arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

诊断智能体编排的视觉-语言-动作技能组合中的语义切换失败

Diagnosing Semantic Handoff Failures in Agent-Orchestrated Vision-Language-Action Skill Composition

Ke Rui, Yushen Zuo, Jiawei Wang, Haoran Jia, Jinming Ma, Weitao Zhou, Minglei Li

arXiv 2607.06256首次发表:更新:

发表机构

SimpleAI(SimpleAI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究智能体编排的视觉-语言-动作技能组合中的语义切换失败问题,通过智能体编排执行框架,调用基于π0.5的技能检查点,在两种初始状态分布下评估,发现单技能与长期任务成功率差距大,为未来技能库改进提供诊断依据。

AI 中文摘要

长期的家庭任务要求机器人组合多种语言条件技能,但连续技能之间的界限很少明确。一项技能可能满足自身后置条件,却使机器人、物体或摄像头视图处于一种状态,导致下一项技能无法可靠启动。我们通过智能体编排的视觉-语言-动作执行框架在BEHAVIOR-1K中研究此语义切换问题。该框架调用基于π0.5的技能检查点,为每个技能分配类型化参数和步骤预算,并使用多视图视觉-语言模型验证来决定执行是继续、重试还是重新规划。为区分孤立技能能力和长期组合鲁棒性,我们在两种初始状态分布下评估相同检查点。选定的导航、抓取、放置和开门技能在人工审核验证下从干净快照中成功率达77%-100%,但组合执行仍常因链式状态停滞。执行跟踪将这些失败归因于下一项技能准备情况、目标基础和低级控制执行,揭示了单技能成功与可靠长期任务完成之间的巨大差距。这些发现将几乎为零的端到端任务成功率转化为可操作的诊断,表明未来的VLA技能库必须学习对干净演示系统性低估的混乱链式状态分布的鲁棒性。

英文摘要

Long-horizon household tasks require robots to compose many language-conditioned skills, yet the boundary between consecutive skills is rarely explicit. A skill may satisfy its own postcondition while leaving the robot, objects, or camera views in a state from which the next skill cannot reliably start. We study this semantic handoff problem in BEHAVIOR-1K through an agent-orchestrated vision-language-action execution harness. The harness invokes $π_{0.5}$-based skill checkpoints trained from cleaned BEHAVIOR-1K demonstrations, assigns each skill typed arguments and a step budget, and uses multi-view vision-language model verification to decide whether execution should advance, retry, or replan. To separate isolated skill competence from long-horizon compositional robustness, we evaluate the same checkpoints under two initial-state distributions: clean skill-boundary snapshots and chained terminal states produced by previous skills. Selected navigation, grasping, placement, and door-opening skills achieve 77--100% success from clean snapshots under human-reviewed verification, yet composed rollouts still frequently stall from chained states. The resulting traces attribute failures to next-skill readiness, target grounding, and control execution, turning nearzero task success into actionable diagnostics for what VLA skill libraries must learn next: robustness to the messy chained-state distribution that clean demonstrations underrepresent.

Journal refRSS SemRob Workshop 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑