发表机构
Xi'an Jiaotong University(西安交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对潜在世界模型正则化中EP目标函数对尾部样本校正梯度不足的问题,提出QQWorld方法,通过分位数-分位数匹配目标函数及跨批次QQ技术,提升了LeWM在四个控制环境中的规划成功率与高斯对齐效果。
AI 中文摘要
潜在世界模型通过在紧凑表示空间中预测未来状态实现高效规划,但其性能高度依赖所学潜在分布的质量。LeWorldModel(LeWM)使用Epps-Pulley(EP)目标函数将其潜在变量正则化为各向同性高斯分布。我们发现,对于孤立的尾部样本,EP的校正梯度会迅速消失,导致重尾偏差得不到充分控制。为解决这一局限,我们提出QQWorld,用分位数-分位数匹配目标函数取代EP,该函数可直接将投影后的潜在样本与秩匹配的高斯分位数对齐,从而在尾部保持有效的校正梯度。我们进一步开发了跨批次QQ,利用之前批次的分离样本扩大有效排序池,并分析其偏差-方差权衡。在四个控制环境中,QQWorld可有效提高LeWM的平均规划成功率,同时始终实现更好的高斯对齐和更薄的潜在尾部。
英文摘要
Latent world models enable efficient planning by predicting future states in a compact representation space, but their performance depends critically on the quality of the learned latent distribution. LeWorldModel (LeWM) regularizes its latents toward an isotropic Gaussian using the Epps-Pulley (EP) objective. We show that the corrective gradients of EP rapidly vanish for isolated tail samples, leaving heavy-tailed deviations insufficiently controlled. To address this limitation, we propose QQWorld, which replaces EP with a quantile-quantile matching objective that directly aligns projected latent samples with rank-matched Gaussian quantiles, thereby maintaining effective corrective gradients in the tails. We further develop cross-batch QQ, which enlarges the effective ranking pool using detached samples from previous batches, and characterize its bias-variance trade-off. Across four control environments, QQWorld effectively improves the average planning success rate of LeWM, while consistently yielding better Gaussian alignment and thinner latent tails.