发表机构
The Chinese University of Hong Kong, Shenzhen; Georgetown University; INSAIT, Sofia University “St. Kliment Ohridski”(香港中文大学(深圳); 乔治城大学; 索菲亚大学“圣·克利门特·奥赫里德斯基”分校INSAIT)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究构建经验敏感型游戏学习框架,通过交互式游戏及相关指标分析人类与自进化语言智能体的游戏决策行为,发现人类会稳定转向全局策略,而自进化智能体的行为转变仍有局限。
AI 中文摘要
大型语言模型智能体越来越多地通过游戏进行评估,但大多数基准测试强调最终结果而非玩家如何从重复交互中学习。我们研究经验敏感型游戏学习:游戏体验如何改变人类和语言智能体的决策行为。我们将经验敏感型游戏学习构建为分析重复游戏中行为变化的框架,而非仅关注最终得分或胜率。我们引入了一套具有可复用策略结构的交互式游戏,以及跨游戏的从贪心到全局的指标和特定游戏的行为诊断方法,这些可从行动轨迹中观测到经验驱动的变化。我们还收集了人类玩家的重复游戏轨迹,并在同一行为指标空间中评估了近期的自进化语言智能体。结果显示,人类玩家表现出可解释且相对稳定的转变,即从局部贪心启发式转向更全局的策略决策。相比之下,当前的自进化智能体常表现出嘈杂且短暂的收益,表明现有的自进化方法在将游戏体验转化为决策行为的持久变化方面仍存在局限。
英文摘要
Large language model agents are increasingly evaluated through games, but most benchmarks emphasize final outcomes rather than how players learn from repeated interaction. We study experience-sensitive game learning: how gameplay experience changes the decision-making behavior of humans and language agents. We formulate experience-sensitive game learning as a framework for analyzing behavioral change across repeated gameplay, rather than only final score or win rate. We introduce a suite of interactive games with reusable strategic structure, together with cross-game greedy-to-global metrics and game-specific behavioral diagnostics that make experience-driven change observable from action traces. We also collect repeated-game trajectories from human players and evaluate recent self-evolving language agents in the same behavioral metric space. Our results show that human players exhibit interpretable and relatively stable shifts from locally greedy heuristics toward more global strategic decisions. In contrast, current self-evolving agents often show noisy and transient gains, suggesting that existing self-evolution methods remain limited in converting gameplay experience into durable changes in decision-making behavior.