arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

MGPO:流形引导的扩散对齐用于任务感知数据集蒸馏

MGPO: Manifold-Guided Diffusion Alignment for Task-Aware Dataset Distillation

Yunyi Chen, Chenru Wang, Xinyi Ye, Zexin Zheng, Chi Zhang

arXiv 2610.05252首次发表:更新:

发表机构

AGI Lab, Westlake University; Eindhoven University of Technology; Guangdong University of Finance and Economics(西湖大学 AGI 实验室; 埃因霍温理工大学; 广东财经大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对扩散数据集蒸馏中密度与判别目标不匹配及几何覆盖缺失问题,提出MGPO,以多目标强化学习结合像素判别与MST几何奖励实现双空间对齐,显著提升下游任务性能。

AI 中文摘要

基于扩散的数据集蒸馏(DD)存在一个根本性的目标不匹配问题:似然驱动的扩散模型优先考虑密度近似,而非下游任务所需的判别性决策边界。除了语义不匹配外,仅依赖密度还会导致几何覆盖损失,即生成的样本坍缩到少数高密度模式中,无法覆盖流形的结构多样性。我们提出了流形引导的策略优化(MGPO),将数据集蒸馏重新表述为一个多目标强化学习问题,并通过像素空间判别奖励和由类内最小生成树(MST)引导的潜在空间几何奖励实现双空间对齐。判别奖励强制类别可分性,而基于MST的几何奖励鼓励生成的潜在变量覆盖每个类别的稀疏几何骨架,从而共同解决这两种失败模式。我们进一步提供了理想化分析,以论证基于MST的奖励的合理性,包括与数据集大小无关的Hausdorff近似界和子采样界。奖励模块化设计通过替换冻结的任务奖励模型,扩展到目标检测和分割等结构化任务。大量实验表明,MGPO持续优于现有方法,包括在低预算设置下分割任务上+8.0%的mIoU提升。

英文摘要

Diffusion-based dataset distillation (DD) suffers from a fundamental objective mismatch: likelihood-driven diffusion models prioritize density approximation over the discriminative decision boundaries required for downstream tasks. Beyond semantic mismatch, relying solely on density also leads to geometric coverage loss, where generated samples collapse into a few high-density modes and fail to cover the manifold's structural diversity. We propose Manifold-Guided Policy Optimization (MGPO), which reformulates DD as a multi-objective reinforcement learning problem and achieves Dual-Space Alignment via a pixel-space discriminative reward and a latent-space geometric reward guided by a class-wise Minimum Spanning Tree (MST). The discriminative reward enforces class separability, while the MST-based geometric reward encourages generated latents to cover a sparse geometric skeleton of each class, jointly addressing both failure modes. We further provide an idealized analysis that motivates the MST-based reward, including a Hausdorff approximation bound and a subsampling bound independent of the dataset size. The reward-modular design extends to structured tasks such as object detection and segmentation by substituting the frozen task reward model. Extensive experiments show MGPO consistently outperforms existing methods, including a +8.0% mIoU gain on segmentation under low-budget settings.

Comments32 pages

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑