arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.29577cs.AIcs.LG

DungeonBench:龙与地下城战斗中规则丰富的战术推理基准

DungeonBench: A Benchmark for Rules-Rich Tactical Reasoning in Dungeons & Dragons Combat

Ismayil Ismayilov, Atakan Kara, Kaan Oktay

首次发表
浏览论文内容

中文总结 AI 辅助

本研究推出DungeonBench基准,用于测试龙与地下城战斗中的规则丰富战术推理,含单次遭遇和关联遭遇两个赛道,评估发现前沿语言模型策略在关联遭遇中存在资源管理等问题

中文摘要 AI 辅助

游戏和模拟器通过将决策转化为可衡量的结果,成为有价值的基准,但许多现有基准套件对规则丰富的战术推理测试不足,战术推理是指当几何、时间、资源、目标和规则交互同时起作用时做出正确选择的能力。我们推出DungeonBench,这是一个针对龙与地下城(Dungeons & Dragons)战斗的战术推理基准,旨在涵盖2014年系统参考文档中绝大多数与战斗相关、可由模拟器解析的内容,同时保留简化战斗模拟器常抽象掉的机制。在每一步,DungeonBench会提供完整的战术观测、待决策项,以及包含移动、攻击、法术、反应、目标、准备和稀缺资源的可执行选项索引列表。任务是评估合法选择,其后果取决于行动经济、生物特性、战场几何、时间窗口和后续遭遇。DungeonBench有两个赛道:遭遇(Encounter)赛道评估单次战斗中的局部战术玩法,每日(Day)赛道则通过持久生命值、法术位、消耗品、准备和短休时间将各遭遇关联起来,迫使策略在即时战术优势与未来生存能力之间权衡。同一引擎生成的决策流支持启发式控制器、语言模型策略、学习到的选项排序器和掩码动作强化学习智能体。我们在该共享决策流上评估前沿语言模型策略,结果显示完整的战术观测并未使基准饱和:前沿策略常能赢得直接遭遇,但关联遭遇的每日任务暴露出资源预算、休息时间安排和规则感知战术纪律方面的缺陷。

英文摘要

Games and simulators make valuable benchmarks by turning decisions into measurable outcomes, but many current suites under-test rules-rich tactical reasoning: the ability to choose well when geometry, timing, resources, objectives, and rule interactions all matter at once. We introduce DungeonBench, a benchmark for tactical reasoning in Dungeons & Dragons combat, built to cover the vast majority of combat-relevant 2014 System Reference Document content whose effects can be resolved by the simulator while retaining mechanics that simplified combat simulators often abstract away. At each step, DungeonBench exposes a complete tactical observation, a pending decision, and an indexed list of executable options spanning movement, attacks, spells, reactions, objectives, preparation, and scarce resources. The task is to value legal choices whose consequences depend on action economy, creature traits, battlefield geometry, timing windows, and future encounters. DungeonBench has two tracks: Encounter, which evaluates local tactical play in single fights, and Day, which links encounters through persistent hit points, spell slots, consumables, preparation, and short-rest timing, forcing policies to trade off immediate tactical advantage against future survivability. The same engine-generated decision stream supports heuristic controllers, language-model policies, learned option rankers, and masked-action reinforcement-learning agents. We evaluate frontier language-model policies on this shared decision stream. Results show that full tactical observations do not saturate the benchmark: frontier policies often win direct encounters, but linked encounter days expose failures in resource budgeting, rest timing, and rule-aware tactical discipline.

↑