智能体 harness:面向机器人自主的大语言模型驱动验证层
Agentic Harnesses: LLM-Driven Verification Layers for Robot Autonomy
浏览论文内容
中文总结 AI 辅助
针对机器人规划模型的安全与伦理风险,提出LLM驱动的验证层作为中间件管控计划,实现近85%的类别准确率、97%的对抗性攻击遏制率,为机器人自主提供可靠保障。
中文摘要 AI 辅助
先进人工智能工具的进展推动了机器人自主领域的研究,但这类系统的开发大多聚焦于执行环节,而非验证规划模型所提动作的可行性。与通用大语言模型(LLM)类似,机器人规划模型存在诸多风险:受用户指定目标的偏向,可能提出不符合科学伦理的动作;因无法“记住”先前的安全风险而存在安全隐患;还可能遭受自主生态系统的对抗性攻击。我们提出一种位于规划与执行之间的大语言模型驱动验证层,用于评估动作的可允许性。我们的“大语言模型作为评判者”集成体结合了各模型的思维链推理,并综合这些专家评判输出,模仿了混合专家与自一致性方法的结合。该层作为中间件,在计划从服务器规划模块到达MCP服务器、进而到达机器人底层控制前对计划进行管控:计划可被批准、因需重新制定而被拒绝,或升级至人工审核。通过该系统,我们在接受/升级/拒绝类别中实现了近85%的准确率,对对抗性攻击的遏制率达97%,接受与拒绝任务间的误差可忽略不计,误差主要出现在升级边界处。
英文摘要
Advances in advanced artificial intelligence tools have sparked research in robot autonomy, but the development of such systems has largely focused on execution rather than verifying the feasibility actions planning models propose. Like general-purpose LLMs, robotics planning models carry risks: biased toward user-specified goals, they may suggest actions misaligned with scientific ethics, they may be unsafe due to an inability to "remember" prior safety risks, or they may be vulnerable to adversarial attacks on the autonomy ecosystem. We propose a LLM-driven verification layer between planning and execution to evaluate action permissibility. Our LLM-as-a-Judge ensemble combines chain-of-thought reasoning across models and synthesizes those expert judge outputs, mirroring a combination of a mixture of experts and self-consistency approach. This layer serves as middleware, gating plans from the server's planning module before they reach the MCP server and therefore the robot's low-level controls: plans are approved, rejected for reformulation, or escalated for human review. With this system, we achieve near 85% precision across accept/escalate/reject categories 97% containment of adversarial attacks, with negligible errors between accepting and rejecting tasks, and errors mostly manifesting at the escalate boundary.
发表机构
- Carnegie Mellon University(卡内基梅隆大学)
- Pacific Northwest National Laboratory(太平洋西北国家实验室)
机构由 AI 辅助整理,请以论文原文为准。