VERA:面向约束多智能体控制的可验证可行性表示与反事实信用
VERA: Verifiable Feasibility Representations with Counterfactual Credit for Constrained Multi-Agent Control
浏览论文内容
中文总结 AI 辅助
VERA提出集中训练分散执行的框架,通过可验证可行性表示和反事实信用,在约束多智能体控制中显著提升成功率并降低违规率。
中文摘要 AI 辅助
约束多智能体控制需要的不仅仅是预测奖励性动作:随着接触窗口、共享容量和截止日期的变化,一个动作可能不再可执行。我们引入VERA,一个集中训练、分散执行的框架,将可行性估计与信用分配分离。每个参与者预测一个五维可验证可行性表示(VFR)。在提出动作后,仅在训练期间可用的精确动作条件边际监督该表示,而反事实组相对优势(CGRA)对候选表示-动作对进行排序。执行使用一次参与者传递且无特权状态。在动态空天地一体化网络(SAGIN)中,VERA获得55.33%±3.60%的成功率,覆盖率违规为0.45%±0.81%,与特权掩码参考相差1.33个百分点。在十个配对种子上匹配奖励后,VERA相比最强基线将成功率提高8.74个百分点(p=0.023),并将违规率降低52.19个百分点(p=5.7e-8)。一个十种子4×2因子实验将14.16-16.48个百分点的增益归因于CGRA,涵盖手工设计、学习、随机和潜在表示;在七个未见拓扑上的评估保持了相对于多智能体近端策略优化的24.33-30.02个百分点优势。从10到40个用户,成功率保持在50.1-53.8%,VFR仅增加0.026毫秒到中央处理器(CPU)参与者步骤。跨域测试进一步确定了主导条件:当候选分数尊重共享约束时,反事实信用成功,而在不兼容的奖励几何下失败。这些结果确立了动作条件可行性作为可审计的训练接口,以及反事实信用作为几何依赖的优化机制。
英文摘要
Constrained multi-agent control requires more than predicting rewarding actions: an action can cease to be executable as contact windows, shared capacity, and deadlines change. We introduce VERA, a centralized-training, decentralized-execution framework that separates feasibility estimation from credit assignment. Each actor predicts a five-dimensional verifiable feasibility representation (VFR). After an action is proposed, exact action-conditioned margins available only during training supervise that representation, while a counterfactual group-relative advantage (CGRA) ranks candidate representation-action pairs. Execution uses one actor pass and no privileged state. In a dynamic space-air-ground integrated network (SAGIN), VERA obtains 55.33% +/- 3.60% success with 0.45% +/- 0.81% coverage violation, within 1.33 percentage points of a privileged-mask reference. With rewards matched over ten paired seeds, VERA improves success over the strongest baseline by 8.74 percentage points (p=0.023) and reduces violation by 52.19 percentage points (p=5.7e-8). A ten-seed 4-by-2 factorial attributes a 14.16-16.48 percentage-point gain to CGRA across handcrafted, learned, random, and latent representations; evaluation on seven unseen topologies preserves a 24.33-30.02 percentage-point advantage over multi-agent proximal policy optimization. From 10 to 40 users, success remains 50.1-53.8%, and VFR adds only 0.026 ms to a central processing unit (CPU) actor step. Cross-domain tests further identify the governing condition: counterfactual credit succeeds when candidate scores respect shared constraints and fails under incompatible reward geometries. These results establish action-conditioned feasibility as an auditable training interface and counterfactual credit as a geometry-dependent optimization mechanism.
发表机构
- Harbin Institute of Technology(哈尔滨工业大学)
- The Chinese University of Hong Kong(香港中文大学)
机构由 AI 辅助整理,请以论文原文为准。