arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

iAm.md:面向未知开放词汇域的智能体内省式机器人技能自我评估

iAm.md: Robot Skill Self-Assessment through Agentic Introspection for Unknown Open-Vocabulary Domains

Vincenzo Guarino, Emanuele Musumeci, Vincenzo Suriani, Daniele Nardi

arXiv 2610.10962首次发表:更新:

发表机构

Sapienza University of Rome(罗马第一大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出iAm.md框架,通过开放词汇语义映射实现智能体内省,在模拟TIAGo机器人任务上验证其可支持技能自我评估与任务泛化,解决机器人自主行为的接地失败问题。

AI 中文摘要

基于大语言模型泛化能力的智能体AI拥有广泛的潜在应用,包括具身任务规划。例如,基于基础模型的具身智能体可在自主机器人场景中生成合理的规划。由于上下文窗口有限或下一个词元预测表述中的幻觉现象,生成的行为可能未确认部署的机器人和观测环境是否实际支持请求的操作,我们将此称为“接地失败”。得益于基础模型推理能力的近期提升,自主机器人行为生成问题可被表述为代码生成问题。我们提出iAm.md,这是一个Markdown标准与生成框架,允许将该过程锚定在互补形式的部署证据中。通过开放词汇语义映射,我们结合局部视觉-语言检测与物体分割,并将其关联到该中间标准化表示中的持久物体记录,从而实现智能体内省。我们在模拟TIAGo机器人的导航与操作任务上研究了该新技术,表明该标准化表示可共同支持技能自我评估与可执行任务泛化。

英文摘要

Agentic AI based on Large Language Model generalization capabilities offers a wide range of potential applications, including planning for embodied tasks. For example, embodied agents based on Foundation models can generate plausible plans in autonomous robotics scenarios. Due to limited context windows or hallucinatory phenomena in the next-token prediction formulation, behaviors may be generated without establishing whether the deployed robot and the observed environment actually support the requested operation, in what we call a "grounding failure". Thanks to the recent improvements in reasoning capabilities of foundation models, autonomous robot behavior generation problem can be formulated as a code generation problem. We present iAm.md, a Markdown standard and generation framework, that allows anchoring this process in complementary forms of deployment evidence. Through open-vocabulary semantic mapping, we combine local vision-language detections and object segmentation and refer them to persistent object records in this intermediate standardized representation, allowing agentic introspection. We then study this new technique on a simulated TIAGo, on navigation-and-manipulation tasks, showing how this standardized representation jointly supports skill self-assessment and executable task generalization.

Comments7 pages, 2 figures, 1 table. Accepted at the 13th Italian Workshop on Artificial Intelligence and Robotics (AIRO 2026). Project page: https://yurimachine.github.io/iAm.md/

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑