发表机构
University of Edinburgh; University of Oklahoma; Imperial College London; University of Michigan; University of Southern California; Tencent; Beijing Normal University; Beijing Normal–Hong Kong Baptist University(爱丁堡大学; 俄克拉荷马大学; 伦敦帝国学院; 密歇根大学; 南加州大学; 腾讯; 北京师范大学; 北京师范大学-香港浸会大学联合国际学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对多智能体环境中对手策略持续进化导致静态技能修改方法失效的问题,提出OASE方法,通过历史快照锚定的配对比较选择有益技能修改,在两类场景中实现更低均衡距离与更少无效策略变更。
AI 中文摘要
通过交互学习调整策略是构建更通用、自主的大语言模型(LLM)智能体的关键一步。现有方法通常通过修改技能库实现行为适应,但在多智能体环境中,对手可能同时更新自身策略,导致环境持续进化。将为静态环境设计的技能修改方法应用于此场景,相当于基于过时参考进行更新。为应对这一挑战,我们提出OASE(Opponent-Aware Selective Evolution,对手感知的选择性进化),该方法可在动态多智能体环境中识别并采用真正有益的技能修改。具体而言,OASE在以对手策略历史快照为锚点的相同条件下,对候选技能与现有技能进行配对比较,仅当候选技能的估计收益增益超过接受阈值时才采用它。我们在两种决策场景中评估OASE:第一价格拍卖和私人成本古诺竞争。实验结果表明,与Reflexion风格基线相比,OASE在两种环境中均实现了更低的最终均衡距离,同时接受的技能修改数量显著更少,从而抑制了缺乏足够收益支持的策略变更。因此,OASE用基于证据锚定的选择取代盲目更新,使智能体即使在对手持续进化的情况下也能稳定高效地适应。
英文摘要
Learning to adapt strategies through interaction is a key step toward more general and autonomous LLM agents. Existing approaches typically achieve behavioral adaptation by revising skill libraries. However, in multi-agent environments, opponents may simultaneously update their strategies, causing the environment itself to evolve continuously. Applying skill-revision methods designed for static environments in such settings therefore amounts to updating against an obsolete reference. To address this challenge, we introduce OASE (Opponent-Aware Selective Evolution), which identifies and adopts genuinely beneficial skill revisions in dynamic multi-agent environments. Specifically, OASE conducts paired comparisons between a candidate skill and the incumbent under identical conditions anchored by historical snapshots of opponent strategies, and adopts the candidate only when its estimated payoff gain exceeds an acceptance threshold. We evaluate OASE in two decision-making scenarios: first-price auctions and private-cost Cournot competition. Experimental results show that, compared with a Reflexion-style baseline, OASE achieves a lower final equilibrium distance in both environments while accepting substantially fewer skill revisions, thereby suppressing strategy changes that lack sufficient payoff support. OASE therefore replaces blind updating with evidence-anchored selection, allowing agents to adapt stably and efficiently even as opponents continuously evolve.