对抗性行动方案生成:用于COA匹配与COA生成的博弈论多智能体算法
Adversarial Course-of-Action Generation: Game-Theoretic Multi-Agent Algorithms for COA matching & COA generation
另 1 家 · 查看机构详情
- Anote AI
- Cornell University(康奈尔大学)
- CUNY(纽约城市大学)
- Stevens Institute of Technology(史蒂文斯理工学院)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
针对COA生成问题,提出COA-Bench离线基准,通过博弈论多智能体算法实现COA匹配与生成,实验表明最佳响应策略和两阶段委员会能有效降低BLUE优势,提升对抗评估可靠性。
中文摘要 AI 辅助
行动方案(COA)生成是一个分布式规划问题:系统必须提出结构化的候选行动,针对对抗性响应进行评估,并呈现那些在不断变化条件下仍保持战术连贯性的选项。我们提出了COA-Bench,一个用于通过自我对弈比较COA生成策略的小型离线基准测试和可复现性工件。遵循BattleCOA术语,我们将COA匹配保留用于资产-效果匹配,将COA生成保留用于行动方案生成;当前工件并未直接实现任何决策函数。相反,它将COA表示为带有条件分支的类型化行动链,分配一个合成的COA质量分数,使用BLUE对RED的优势分数和纳什间隙距离来比较对立的COA,并使用受FM 3-0启发的启发式评分标准对条令一致性进行评分。在跨越五种作战模板的50个合成场景中,一种采样最佳响应策略(抽取8个RED候选)将BLUE优势从0.516降至0.485,并将BLUE兵棋胜率从0.920降至0.820;一个两阶段多智能体委员会(包含5个BLUE提议智能体、RED团队裁决和基于批评的修订)获得了0.509的BLUE优势和0.820的BLUE胜率。我们还识别并修复了一个基准设计问题,即场景框架被存储为元数据,但对生成的COA内容没有影响。COA-Bench不是一个作战管理系统,也不使用任何真实、机密、专有或人类受试者数据。其贡献在于一个可检查的评估框架、初步基准证据,以及构建可审计的智能体规划工件的经验教训。
英文摘要
Course-of-action (COA) generation is a distributed planning problem: a system must propose structured candidate actions, evaluate them against an adversarial response, and surface options that remain tactically coherent under changing conditions. We present COA-Bench, a small offline benchmark and reproducibility artifact for comparing COA generation policies through self-play. Following the BattleCOA terminology, we reserve COA matching for asset-effect matching and COA generation for course-of-action generation; the present artifact does not implement either DecisionFunction directly. Instead, it represents COAs as typed action chains with conditional branches, assigns a synthetic COA quality score, compares opposing COAs with a BLUE-vs-RED advantage score and Nash-gap distance, and scores doctrinal coherence with an FM 3-0-inspired heuristic rubric. Across 50 synthetic scenarios spanning five operational templates, a sampled best-response policy that draws eight RED candidates reduces BLUE advantage from .516 to .485 and BLUE wargame win rate from .920 to .820; a two-stage multi-agent council with five BLUE proposer agents, RED-team adjudication, and critique-driven revision obtains .509 BLUE advantage and .820 BLUE win rate. We also identify and fix a benchmark-design issue in which scenario framing was stored as metadata but had no effect on generated COA content. COA-Bench is not an operational battle-management system and uses no real, classified, proprietary, or human-subject data. The contribution is an inspectable evaluation harness, preliminary benchmark evidence, and lessons for building auditable agentic planning artifacts.