arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.24741cs.DCcs.LG

用于动态熵最优传输的经认证的并行时间Sinkhorn算法

DrainSinkhorn: Safe Elimination for Batched Entropic Optimal Transport

Xinyang Wen

首次发表
浏览论文内容

中文总结 AI 辅助

研究动态应用中熵最优传输问题,提出并行时间执行器TemporalSinkhorn,通过批处理、证书、遗忘率等机制优化,在多GPU及不同场景下实验,显著提升运行速度且保证输出精度,为相关领域提供高效解决方案。

中文摘要 AI 辅助

动态应用,包括最优传输流匹配,需反复解决相关熵最优传输问题,但传统分布式Sinkhorn过程是顺序执行且每次迭代后同步。我们提出了TemporalSinkhorn,一种并行时间执行器,它对未来候选者及其修复进行批处理,且不使输出精度具有推测性。一个居中的、按行分片的证书仅接受确定性的安全前缀。其余候选者共享打包的Sinkhorn更新;在线投影遗忘率设置审核里程碑,而后验残差检查从每个深度低估中恢复。预测可改变工作安排,但不会授权不准确的输出。在4个A100 GPU上,n = 2048时60次运行、五种子网格的实验表明,相对于在五个统计解析的区域单元中对每个打包迭代进行审核,遗忘引导的里程碑将运行时间减少了1.15倍至1.47倍。在六个合成流上,相对于顺序软c变换热启动,并行时间执行快1.42倍至3.55倍,且无边际容差违规。在流匹配小批量流上,n = 2048时并行时间执行比顺序执行快3.054倍至3.632倍,无容差违规。在RTX 4060笔记本GPU上的单独固定内核测试给出了4.315倍的几何平均加速比。这些是互补的部署研究,而非受控硬件比较。端到端流匹配集成、优化求解器比较和多节点验证仍未完成。

英文摘要

Fast entropic optimal transport backends reduce the cost of each Sinkhorn update, but static batches still run at full width until the slowest problem finishes. We introduce DrainSinkhorn, a verifier-gated active-packing layer for batches of independent Sinkhorn problems. It combines candidate-axis packing, a Sinkhorn-specific one-sided screen, verifier-gated retirement under the backend's configured two-sided residual check, and physical compaction of all candidate-indexed state. The EOT objective, per-instance Sinkhorn map, and stopping rule are unchanged; later kernels run only on unfinished problems. We characterize the removable work exactly. If completion depths differ within a packed window, active execution removes the padding between the static batch rectangle and the observed survival curve. A quotient nonlinear Perron-Frobenius analysis gives a local explanation for these finite-tolerance depth differences: convergence depends on the full modal spectrum and proposal alignment, not only on the slowest mode. DrainSinkhorn achieves state-of-the-art execution performance on the tested heterogeneous batched-EOT workloads within matched backend families. The complete Flash-backed OT path is 4.110x faster on MetroPT-3, 3.798x faster on ImageNet-32 feature couplings, and 1.250-1.270x faster across a five-tolerance Packer19 sweep. Independent implementations reach 2.600x on ImageNet-32 with OTT-JAX, 3.174x on A2D2 LiDAR with PyKeOps, and 1.415x on large ImageNet-32 PyKeOps couplings. End-to-end speedups remain 4.074x on MetroPT-3 and 2.786x on ImageNet-32 feature-space OT flow matching, with all reported residual, consumer-output, and training-quality checks passing.

补充信息

↑