发表机构
College of Information Engineering, Hunan Applied Technology University(湖南应用技术学院信息工程学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对并行状态熵优化中的冗余探索问题,提出MCC-PGPSE方法,通过分配边际覆盖率信用重新分配辅助奖励,在多类基准上提升了并行策略的互补覆盖效果。
AI 中文摘要
用于并行状态熵最大化的策略梯度算法(PGPSE)通过在同一环境的复本中训练独立参数化的策略,扩展了状态空间覆盖范围。然而,其聚合的团队熵分数仅衡量集体探索情况,无法识别出能贡献非冗余覆盖的策略。我们为PGPSE引入了边际覆盖率信用(MCC-PGPSE),该方法结合了留一策略覆盖与状态所有者专业化来估计策略特定的信用。MCC-PGPSE保留了PGPSE的聚合目标,并根据这些信用重新分配非负的辅助内在奖励,同时不改变其总质量。这种重新分配旨在抑制冗余访问并促进互补覆盖。我们在受控环境、7个公开离散状态基准,以及原始PGPSE协议中的代表性Room和Maze设置中评估了MCC-PGPSE。在所有测试设置中,MCC-PGPSE在归一化团队状态熵和状态支持度上,相对于熵基线产生了正的最终窗口增益。受控任务比较和固定套件公开聚合结果具有统计学显著性,而原始协议的五种子比较结果方向一致。消融实验和信用对齐控制实验表明,大部分增益来自留一策略覆盖,而非非均匀加权、不匹配的信用或单纯的神经新颖性。这些结果支持基于贡献的辅助奖励分配作为一种可解释的方法,用于改善离散状态空间中并行策略间的互补覆盖。
英文摘要
Policy Gradient for Parallel State Entropy maximization (PGPSE) expands state-space coverage by training independently parameterized policies in replicated copies of the same environment. However, its pooled team-entropy score measures only collective exploration and cannot identify policies that contribute non-redundant coverage. We introduce Marginal Coverage Credit for PGPSE (MCC-PGPSE), which combines leave-one-policy-out coverage with state-owner specialization to estimate policy-specific credit. MCC-PGPSE preserves PGPSE's pooled objective and redistributes non-negative auxiliary intrinsic rewards according to these credits without changing their total mass. This redistribution is designed to discourage redundant visitation and promote complementary coverage. We evaluated MCC-PGPSE in controlled environments, seven public discrete-state benchmarks, and representative Room and Maze settings from the original PGPSE protocol. Across all tested settings, MCC-PGPSE produced positive final window gains in normalized team state entropy and state support over the Entropy baseline. Controlled-task comparisons and the fixed-suite public aggregate were significant, whereas five-seed original-protocol comparisons were directionally consistent. Ablations and credit alignment controls indicate that most gains arise from leave-one-policy-out coverage rather than non-uniform weighting, mismatched credit, or neural novelty alone. These results support contribution-conditioned auxiliary reward allocation as an interpretable approach to improving complementary coverage among parallel policies in discrete state spaces.