AI 中文总结
本文提出从风险评分到风险分配的密度驱动框架,识别拥挤悖论并利用密度信号,通过QUBO优化实现多智能体监控中的风险-多样性权衡,显著提升性能。
AI 中文摘要
多智能体系统中的风险监控通常建立在逐状态基元之上,该基元独立地对每个状态进行评分并选择前K个状态。在拥挤场景下,当许多智能体共享相同的脆弱性时,这种方法会选择冗余的警报,其风险是联合相关的,我们将这种模式描述为“监控中的羊群效应”。我们提出从风险评分到风险分配的范式转变,并由两项贡献支持。首先,我们识别出拥挤悖论,即P(风险 | x) ∝ p(x)而非1/p(x),因此密度而非异常分数是有效的风险信号;在金融数据上,基于密度的评分在5天/10天/20天崩溃时间范围内达到AUROC ≥ 0.94,而五个异常基线均低于0.80。其次,给定一个密度派生的脆弱性分数,我们将监控重新构建为对相互依赖状态的组合子集选择,并将其映射到具有λ控制的风险-多样性权衡的QUBO目标。由此产生的帕累托前沿包含标准多样子集方法(MMR、k-DPP)作为固定工作点;相对于贪心方法的增益随规模单调增加,从n=15时的+24%到n=200时的+66%;学习到的λ策略达到Oracle网格搜索目标的99.5%;该公式可迁移到交通和多智能体强化学习。相同的QUBO实例无需修改即可在Rigetti超导量子处理单元(通过Amazon Braket的Ankaa-3和Cepheus-1-108Q)上执行,我们将其报告为该公式的兼容性属性,而非在此规模下量子优势的声明。
英文摘要
Risk monitoring in multi-agent systems is commonly built on a per-state primitive that scores each state independently and selects the top K. Under crowding, where many agents share the same fragility, this approach picks redundant alerts whose risks are jointly correlated, a pattern we describe as ``herding in monitoring.'' We propose a paradigm shift from risk scoring to risk allocation, supported by two contributions. First, we identify the Crowding Paradox, namely that P(risk | x) $\propto$ p(x) rather than 1/p(x), so density rather than anomaly score is the operative risk signal; on financial data, density-based scoring reaches AUROC $\geq$ 0.94 at 5d/10d/20d crash horizons, while five anomaly baselines all fall below 0.80. Second, given a density-derived fragility score, we recast monitoring as combinatorial subset selection over interdependent states and map it to a QUBO objective with a $λ$-controlled risk--diversity tradeoff. The resulting Pareto frontier contains standard diverse-subset methods (MMR, k-DPP) as fixed operating points; the gain over greedy grows monotonically with scale, from +24% at n=15 to +66% at n=200; a learned $λ$ policy reaches 99.5% of an oracle grid-search objective; and the formulation transfers to traffic and multi-agent reinforcement learning. The same QUBO instances execute without modification on Rigetti superconducting QPUs (Ankaa-3 and Cepheus-1-108Q via Amazon Braket), which we report as a compatibility property of the formulation rather than a claim of quantum advantage at this scale.
Comments13 pages, 7 figures. Accepted at the ICML 2026 Workshop on New Frontiers in Game-Theoretic Learning (NExT-Game)