arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

FairTest:多智能体强化学习系统的基于搜索的公平性测试

FairTest: Search-Based Fairness Testing for Multi-Agent Reinforcement Learning Systems

Xiaotong Wang, Xuan Xie

arXiv 2609.27309首次发表:更新:

发表机构

Macau University of Science and Technology(澳门科技大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对多智能体强化学习系统公平性测试缺失的问题,提出基于搜索的FairTest方法,结合三适应度引导与优先级排序,在三个环境上检测到最多公平性故障,故障数平均超最强基线221%。

AI 中文摘要

多智能体强化学习(MARL)训练一组共享同一环境并共同学习其策略的智能体。训练旨在最大化团队回报,但高回报并不意味着在每个回合中奖励在智能体之间公平分配。测试是发现深度强化学习故障的既定方法,但很少有方法解决MARL的公平性问题。在这项工作中,我们提出了FairTest,一种基于搜索的测试方法,旨在寻找MARL策略的不公平执行。该设计将搜索引导与测试优先级排序相结合。引导机制使用三个适应度函数对每个候选进行评估:一个衡量已执行运行的公平性,另一个从抽象状态和公平性特征预测公平性,第三个从策略中读取决策不确定性。交叉和变异从观察到的执行中衍生出进一步的候选。优先级排序根据预测的公平性和决策不确定性对候选进行排序,以便运行到达预期出现故障的候选。FairTest在三个环境和两种MARL算法上进行了评估,并给予四个基线相同的预算。与三个基线相比,它检测到最多的公平性故障,具有统计显著性和较大的效应量。故障数量平均超过最强基线的221%,覆盖率平均提高23%。

英文摘要

Multi-agent Reinforcement Learning (MARL) trains a team of agents that share one environment and learn their policies together. Training maximizes the team return, and a high return does not imply that the rewards are shared fairly among the agents in every episode. Testing is an established way to discover the failures of deep reinforcement learning, yet few methods address the fairness of MARL. In this work, we propose FairTest, a search-based testing approach that seeks the unfair executions of a MARL policy. The design combines search guidance with test prioritization. The guidance scores each candidate with three fitness functions. One measures the fairness of the runs already performed, another predicts the fairness from abstract states and fairness features, and the third reads the decision uncertainty from the policy. Crossover and mutation derive further candidates from the observed executions. The prioritization ranks the candidates by the predicted fairness and the decision uncertainty, so that the runs reach the candidates where failures are expected. FairTest is evaluated on three environments and two MARL algorithms, and four baselines are given the same budget. It detects the most fairness failures compared to three baselines with statistical significance and large effect sizes. The failure count exceeds that of the strongest baseline by 221% on average and coverage improves by an average of 23%.

Comments28 pages, 7 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑