arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.06762cs.LG

基于近似最近邻的次二次双模拟度量:覆盖增强保证与可计算的双面证书

Sub-Quadratic Bisimulation Metrics via Approximate Nearest Neighbors: Coverage-Augmented Guarantees and Computable Two-Sided Certificates

Ibne Farabi Shihab, Joyanta Jyoti Mondal

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对MDP提出带证书的次二次双模拟度量方法,通过近似最近邻索引与单调上下界实现,在基准测试中比基线更高效,且在网格世界任务中性能提升28.6%。

中文摘要 AI 辅助

双模拟度量量化马尔可夫决策过程(MDP)中的行为相似性,但其Wasserstein不动点算子需更新每一对状态,产生二次级别的成对计算开销。针对具有有界转移支持和有用低维索引表示的MDP,我们提出一种带证书的次二次方法:近似最近邻索引选择精确受限算子更新的状态对,单调上下界序列在每一轮迭代中包围精确度量。主要分析结果为覆盖增强的任意时间界:仅局部索引质量无法控制全局误差,因为未覆盖的状态对仍保留初始值差距;极限误差最多为max(ρ, εop/(1−γ)),且在精确覆盖备份下,下界序列满足||d̲−d||_∞=ρ,由于ρ依赖未知的精确度量,算法返回可观测的区间宽度,诱导的上下聚类的一致性可证明覆盖聚合的精确恢复。无奖励的下界表明,先覆盖的次二次索引无法消除覆盖项,而单独的自适应下界需要Ω(|S|)次对评估。精确算子实验验证了所有种子运行中的恒等性和包围性,计时实验在廉价和完整Wasserstein备份下均呈现二次与次二次的缩放差异。在分组的|S|=64基准上,精确受限细化在检索覆盖约一半状态对时达到精确度量上限,而独立训练的MICo和DBC基线在所有检索预算下比该上限高22至33倍;Taxi在无信息嵌入下显示证书弃权(不执行),而2500状态网格世界使用一次二次扫描的12.8%计算量,比仅奖励度量提升28.6%。

英文摘要

Bisimulation metrics quantify behavioral similarity in Markov decision processes, but their Wasserstein fixed-point operator updates every state pair and incurs quadratic pairwise work. We give a certificate-carrying sub-quadratic method for MDPs with bounded transition support and a useful low-dimensional indexing representation: an approximate-nearest-neighbor index selects the pairs updated by the exact restricted operator, while monotone lower and upper runs enclose the exact metric at every sweep. The main analytical result is a coverage-augmented anytime bound: local index quality alone cannot control global error, because uncovered pairs retain their initialization gap. The limiting error is at most $\max(ρ,\eop/(1-γ))$, and with exact covered backups the lower arm satisfies $\|\dann-d\|_\infty=ρ$. Because $ρ$ depends on the unknown exact metric, the algorithm returns the observable sandwich width instead; agreement of the induced lower and upper clusterings certifies exact recovery of the covered aggregation. A reward-oblivious lower bound shows sub-quadratic index-first coverage cannot remove the coverage term, while a separate adaptive lower bound requires $Ω(|\Scal|)$ pair evaluations. Exact-operator experiments verify the identity and enclosure in every seeded run, and timing experiments recover quadratic versus sub-quadratic scaling under both cheap and full Wasserstein backups. On the grouped $|\Scal|=64$ benchmark, exact restricted refinement reaches the exact-metric skyline once retrieval covers roughly half of all pairs, while independently trained MICo and DBC baselines stay $22$-$33\times$ above that skyline at every retrieval budget. Taxi shows the certificate abstaining under an uninformative embedding, while a $2500$-state gridworld improves over a reward-only metric by $28.6\%$ using $12.8\%$ of one quadratic sweep.

补充信息

↑