arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

OverForge:通过策略与战术推理实现协作式终身适应

OverForge: Reasoning Through Strategies and Tactics Helps Cooperative Lifelong Adaptation

Oana Madalina Fron, Ojas Shirekar, Chirag Raman

arXiv 2609.39727首次发表:更新:

发表机构

Delft University of Technology(代尔夫特理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

OverForge提出一种免训练的分层架构,将策略推理与战术推理分离,通过元认知模块耦合,在OvercookedV2中实现更高的协作效率、角色保留和适应陌生伙伴,支持终身适应。

AI 中文摘要

协作式语言模型智能体必须在长时间跨度内进行协调,并适应不断变化的环境以及具有不熟悉惯例的合作伙伴,然而现有智能体直接将观察映射到行动,而未将持久协调策略与其战术执行分离。我们提出OverForge,一种免训练的分层架构,将关于角色和分工的策略推理与每个智能体私有的、以合作伙伴为条件的世界模型中的行动战术推理分离开来。一个元认知的前额叶皮层模块通过形成策略-行动分支、使用前向模型想象其后果并在有把握时做出承诺,将这两个层次耦合起来。在OvercookedV2中,OverForge在连通厨房中交付7碗汤,而每个扁平LLM基线仅交付3碗,它保留已商定的角色,并采纳不熟悉合作伙伴提出的角色。消融实验和固定策略探针表明,持久策略指导战术适应,而每个推理层次都对协调有所贡献。记忆重启实验表明,跨情节的合作伙伴知识支持任务性能和合作伙伴预测,将该层次结构与持续适应联系起来。

英文摘要

Cooperative language-model agents must coordinate over long horizons and adapt to changing environments and to partners with unfamiliar conventions, yet existing agents map observations to actions without separating persistent coordination strategies from their tactical execution. We introduce OverForge, a training-free hierarchical architecture that separates strategic reasoning over roles and divisions of labour from tactical reasoning over actions within each agent's private, partner-conditioned world model. A metacognitive Prefrontal Cortex Module couples the two levels by forming strategy-action branches, imagining their consequences with a forward model, and committing when confident. In OvercookedV2, OverForge delivers 7 soups in a connected kitchen versus 3 for each flat LLM baseline, retains agreed roles, and adopts roles proposed by unfamiliar partners. Ablations and a fixed-strategy probe show that persistent strategies guide tactical adaptation while each reasoning level contributes to coordination. Memory restarts show that cross-episode partner knowledge supports task performance and partner prediction, linking the hierarchy to continual adaptation.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑