arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.27365cs.MA

锚定与扰动:通过探索注入实现惰性智能体的补救

Anchor and Perturb: Lazy Agent Remediation by Exploration Injection

Chengxi Zhong, Yongzhe Chang

首次发表
浏览论文内容

中文总结 AI 辅助

AnP框架通过解耦探索注入与流形稳定性,隔离惰性智能体并注入不对称探索脉冲,同时锚定收敛队友,从而挽救崩溃策略并提升胜率至90%。

中文摘要 AI 辅助

锚定与扰动(Anchor and Perturb, AnP)是一种轻量级框架,通过将探索性方差注入与循环流形稳定性解耦,解决多智能体协调失败问题。现有补救策略主要侧重于修改混合网络架构或在整个智能体群体中强制同时探索,这不可避免地会在非单调奖励空间中引发严重的时间差分惩罚。具体而言,AnP隔离表现不佳的惰性智能体,向目标坐标注入不对称的探索脉冲,同时将已收敛的队友锚定在名义上的贪婪利用策略上。经验遥测基准测试表明,AnP成功挽救了崩溃的联合策略(从5%的评估胜率低谷恢复到85%),并促进从次优协调平台期逃脱,持续维持90%的峰值胜率,且无需进行结构性网络修改。

英文摘要

Anchor and Perturb (AnP) is a lightweight framework that resolves multi-agent coordination failures by decoupling exploratory variance injection from recurrent manifold stability. Existing remediation strategies predominantly alter mixing network architectures or enforce simultaneous exploration across the collective, which inevitably precipitates severe temporal-difference penalties in non-monotonic reward spaces. Specifically, AnP isolates underperforming lazy agents and injects an asymmetric exploratory pulse into targeted coordinates whilst anchoring converged teammates to nominal greedy exploitation. Empirical telemetry benchmarks demonstrate that AnP successfully rescues collapsed joint policies (recovering from a 5% evaluation win rate nadir back to 85%) and facilitates escape from suboptimal coordination plateaus, sustaining peak win rates of 90% without requiring structural network modifications.

补充信息

↑