弹性对具身智能体系统的重要性:新指标、系统评估与优化
Resilience Matters for Embodied Agents System: New Metrics, Systematic Evaluation, and Optimization
浏览论文内容
中文总结 AI 辅助
该研究针对具身智能体系统忽略弹性属性的问题,提出弹性评估框架与指标集,经400项家庭任务实验验证其诊断优化效果,揭示弹性特征间的权衡关系。
中文摘要 AI 辅助
具身智能体系统(Embodied Agents System, EAS)正越来越多地部署在开放世界物理领域,其可靠性直接决定部署质量和人机信任。然而,现有评估依赖成功率或安全分数等以结果为中心的指标,这些指标将多样的执行轨迹简化为粗略分数,掩盖了智能体行为背后的动态过程,因此忽略了EAS的一个关键属性——我们将其定义为弹性,该属性反映EAS在扰动下及迭代更新中的恢复、稳定与扩展能力。在开放世界环境中,由于持续存在意外干扰,弹性缺失尤为关键,直接影响EAS部署质量。为解决该问题,我们从弹性工程概念中获取EAS基础的洞见,提出一种可灵活应用于任何EAS的新型弹性评估框架。具体而言,我们定义了首个针对EAS系统的综合弹性指标集,涵盖具身任务执行中的回弹、稳定与优雅扩展性,为EAS弹性分析提供实用基础。我们进一步实现了弹性评估层,将执行过程转化为诊断与优化的评估依据。在10个EAS的400项家庭任务中,我们揭示了结果指标隐藏的过程级差异,包括成功 episodes间的恢复成本差异(ΔC_rec=25.2)、不稳定性增加及任务家族退化。指标引导的优化降低了恢复成本,提高了稳定性和优雅扩展性完成度,展现了弹性评估的诊断效果。我们的结果揭示了弹性特征间的权衡关系,表明构建弹性EAS应根据特定部署需求进行配置。
英文摘要
Embodied Agents System (EAS) are increasingly deployed in open-world physical domains, where reliability directly dictates deployment quality and human-agent trust. However, existing evaluations rely on outcome-centric metrics as success rate or safety scores that collapse diverse execution trajectories into coarse scores, obscuring the dynamic processes underlying agent behavior. Therefore, they ignore a critical property of EAS -- which we define as the Resilience -- that reflects how EASs recover, stabilize, and extend under perturbations and across iterative updates. The lack of resilience is particularly critical in open-world environments due to continuous unexpected disruptions, thus directly affecting the quality of EAS deployment. To address this problem, we gain insight from the resilience-engineering concepts to EAS groundings and propose a novel resilience evaluation framework that can be flexibly applied to any EAS. Specifically, we define the first comprehensive resilience metrics suite for EASs system that exposes Rebound, Stability, and Graceful Extensibility across embodied tasks execution, providing a practical grounding for EAS resilience analysis. We further implement the resilience evaluation layer that transforms execution process into assessments for diagnosis and optimization. Across 400 household tasks with 10 EAS, we reveal the process-level distinction hidden by outcome metrics, including recovery cost differences among successful episodes ($ΔC_{rec}=25.2$), increased instability and task-family degradation. Metrics-guided optimizations reduce recovery cost and increase stability, graceful extensibility completion, showing the diagnostic effect of resilience evaluation. Our results reveal a trade-off among resilience characteristics, suggesting that a resilient EAS construction should be configured according to deployment-specific requirements.
发表机构
- National University of Defense Technology(国防科技大学)
- Nanyang Technological University(南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。