AI 中文总结
研究基于大语言模型的游戏智能体在复杂任务中表现不佳的问题,通过实验验证因果提示增强和多步规划可提升胜率并控制延迟,引入新基准评估,发现较大模型定位更准,融入因果上下文和多步规划能优化性能与速度。
AI 中文摘要
基于大语言模型的游戏智能体在更复杂任务上表现不佳。本文研究这些失败是否与有限的空间推理有关,并评估因果提示增强和多步规划能否在控制响应延迟的同时提高胜率。使用开源Qwen3模型家族,在不同模型规模、推理模式和规划范围内进行实验。引入由三个定制游戏和五个难度级别组成的GVGAI基准来隔离空间导航。通过定位实验和游戏成功研究两种范式评估。结果表明,虽启用思维模式的较大模型能更准确定位,但较小模型坐标匹配性能仍有限。随着游戏级别和布局复杂度增加胜率降低。将因果上下文融入提示往往能提高成功率,启用思维模式和更长规划范围可显著提升性能,多步规划还能进一步减少平均每步响应时间。
英文摘要
LLM-based game agents often perform poorly on more complex tasks. This work examines whether these failures are linked to limited spatial reasoning and evaluates whether causal prompt augmentation and multi-step planning can improve win-rates while managing response latency. Using the open-source Qwen3 model family, we conduct experiments across varying model scales, reasoning modes, and planning horizons. We further introduce a focused GVGAI benchmark consisting of three custom games with five difficulty levels to isolate spatial navigation. The evaluation follows two paradigms: an initial ``positioning experiment'' to test an agent's ability to find its exact coordinates, and a study of game-play success. Our results show that while larger models with an enabled thinking mode identify their positions more accurately, overall performance in coordinate matching remains limited for smaller models. Win rates decrease as game levels and layout complexity increase, validating the benchmark's difficulty scaling. Integrating causal context into the prompts tends to improve the agents' success rates, particularly for bigger models. While enabling thinking mode and longer planning horizons significantly improve performance, multi-step planning further reduces mean per-step response times, offering a practical trade-off between reasoning depth and execution speed.
CommentsTo be published at COG 2026