arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

面向隐藏目标下零样本人机协作的结构化大语言模型推理

Structured LLM Reasoning for Zero-Shot Human--Robot Coordination Under Hidden Goals

Dong Hae Mangalindan, Anand Gokhale, Francesco Bullo, Vaibhav Srivastava

arXiv 2608.04309首次发表:更新:

发表机构

Michigan State University; UC Santa Barbara(密歇根州立大学; 加州大学圣巴巴拉分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对隐藏目标下的零样本人机协作,提出结构化LLM架构,通过Dec-POMDP分解决策,实验显示其交互步骤更少、人类信任度更高,优于无ToM推理的变体和离线训练的多智能体强化学习策略。

AI 中文摘要

我们提出一种结构化大语言模型(LLM)架构,用于在具有私有目标视图的协作构建任务中实现零样本人机协作。在Dec-POMDP公式的引导下,该架构将决策过程分解为五个部分:(i)动作条件化的心智理论(ToM)推理,(ii)分层规划,(iii)对话解释,(iv)动作验证,(v)基于反馈的重新规划。我们将所提方法与未采用ToM推理的 ablation 变体,以及在大量目标对上离线训练的多智能体强化学习策略进行对比。在人类参与者实验中,所提方法所需交互步骤更少,且比两个基线方法产生更高的交互后信任评分。这些结果表明,系统分解团队决策问题、利用LLM作为原本难以处理的推理与规划计算的可处理代理,以及保留物理可行性的常规验证,可同时改善任务协作与人类体验。

英文摘要

We present a structured large-language-model (LLM) architecture for zero-shot human--robot coordination in a cooperative construction task with private goal views. Guided by a Dec-POMDP formulation, the architecture decomposes decision-making into (i) action-conditioned Theory-of-Mind (ToM) inference, (ii) hierarchical planning, (iii) conversation interpretation, (iv) action verification, and (v) feedback-based replanning. We compare the proposed method with an ablation without ToM inference and a multi-agent reinforcement-learning policy trained offline over many goal pairs. In human-participant experiments, the proposed method required fewer interaction steps and yielded higher post-interaction trust ratings than both baselines. These results suggest that systematically decomposing the team decision problem, using LLMs as tractable surrogates for otherwise intractable inference and planning computations, and retaining conventional verification for physical feasibility can improve both task coordination and the human experience.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑