AdvPlan-Bench:结构化计划生成智能体的对抗性评估基准
AdvPlan-Bench: Adversarial Evaluation of Structured Plan-Generation Agents
浏览论文内容
中文总结 AI 辅助
本文提出AdvPlan-Bench,这一结构化计划生成智能体的对抗性评估离线基准,通过多维度指标开展实验验证,为相关研究提供可复现的基准支持。
中文摘要 AI 辅助
结构化计划生成智能体的评估常孤立考察计划本身的质量,但现实中很多规划任务需考量候选计划在其他智能体搜索响应时的表现。本文提出AdvPlan-Bench,这是一个用于结构化计划生成智能体对抗性评估的离线基准。其核心贡献是一套通用评估对象:类型化计划、对抗响应集、选择器诊断及可追踪的候选前沿指标。AdvPlan-Bench将计划表示为带可选分支的类型化动作链,分配合成质量分数,通过BLUE-vs-RED优势和纳什差距诊断对比对立计划,并用透明启发式规则评估定性约束一致性。在涵盖5种规划模板的150个合成场景中,抽取8个响应候选的采样最佳响应策略,相较于单样本响应,使BLUE优势从0.518降至0.486,BLUE胜率从0.900降至0.820;离线LLM策略契约基线的BLUE优势为0.496、胜率0.700,两阶段多智能体委员会的对应值为0.509、0.813。对600条评分记录的三评分者规则敏感性研究显示,评分者间一致性达0.978。需说明的是,AdvPlan-Bench并非可运行规划器,也不提供现实决策质量的证据,它是用于研究对抗性计划评估、响应预算敏感性、候选前沿及多智能体批判与修正轨迹的可复现基准制品。
英文摘要
Structured plan-generation agents are often evaluated as if a plan has quality in isolation, yet many realistic planning tasks require asking how a candidate behaves when another agent can search for responses. We introduce AdvPlan-Bench, an offline benchmark for adversarial evaluation of structured plan-generation agents. The contribution is a general evaluation object: a typed plan, an adversarial response set, selector diagnostics, and traceable candidate-frontier metrics. AdvPlan-Bench represents plans as typed action chains with optional branches, assigns synthetic quality scores, compares opposing plans with BLUE-vs-RED advantage and Nash-gap diagnostics, and evaluates qualitative constraint coherence with a transparent heuristic rubric. In 150 synthetic scenarios spanning five planning templates, a sampled best-response policy that draws eight response candidates reduces BLUE advantage from .518 to .486 and BLUE win rate from .900 to .820 relative to a single-sample response. An offline LLM-policy contract baseline reaches .496 BLUE advantage and .700 BLUE win rate, while a two-stage multi-agent council obtains .509 BLUE advantage and .813 BLUE win rate. A three-rater rubric-sensitivity study over 600 rating records yields .978 inter-rater agreement. AdvPlan-Bench is not an operational planner and provides no evidence about real-world decision quality; it is a reproducible benchmark artifact for studying adversarial plan evaluation, response-budget sensitivity, candidate frontiers, and multi-agent critique-and-revision traces.