arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

贝叶斯伙伴建模支持大语言模型协同的自适应重规划

Bayesian Partner Modelling enables Adaptive Replanning for LLM Coordination

Harsh Goel, Aditya Sai Ellendula, Vaishnav Tadiparthi, Ehsan Moradi Pari, Hossein Nourkhiz Mahjoub, Sandeep P. Chinchali

arXiv 2608.18490首次发表:更新:

AI 中文总结

该研究针对LLM多智能体协作中伙伴策略变化导致的重规划问题,提出BayesBeliefAgent方法,在Overcooked环境中显著缩小信念-行动差距并减少重规划次数。

AI 中文摘要

多智能体大语言模型(LLM)系统常难以与任务中途策略变化的新队友协作,由于智能体执行多步骤或时间扩展技能,在公共证据显示伙伴已改变技能后,它们仍会继续执行过时计划。现有方法要么将伙伴跟踪视为被动上下文(使智能体知晓变化但行动迟缓),要么无差别重规划。我们提出BayesBeliefAgent,它将分层LLM规划器与贝叶斯跟踪模块结合,智能体仅在伙伴行为与推断技能直接矛盾时才中断当前技能,而非持续重规划。除标准奖励外,我们用重规划效率和信念-行动差距(正确估计伙伴的智能体执行非互补技能的决策占总决策的比例)评估性能。在基准Overcooked环境中,矛盾条件控制大幅缩小了该信念-行动差距,且重规划次数比启发式方法少一个数量级。

英文摘要

Multi-agent Large Language Model (LLM) systems often struggle to collaborate with new teammates whose strategies shift mid-task. Because agents execute multi-step or temporally extended skills, they frequently continue executing outdated plans long after public evidence shows that a partner has changed its skill. Existing methods either treat partner tracking as passive context-leaving the agent aware of the shift but slow to act-or replan indiscriminately. We introduce BayesBeliefAgent, which pairs a hierarchical LLM planner with a Bayesian tracking module. Rather than replanning constantly, our agent interrupts its current skill only when a partner's actions directly contradict the inferred skill. Beyond standard reward, we evaluate performance using replanning efficiency and the belief-action gap: the fraction of total decisions where an agent with a correct partner estimate executes a non-complementary skill. Across benchmark Overcooked environments, contradiction-conditioned control drastically narrows this belief-action gap while requiring an order of magnitude fewer replans than heuristic methods

CommentsPreprint

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑