发表机构
Huazhong University of Science and Technology; Ant Group; Chinese Academy of Sciences(华中科技大学; 蚂蚁集团; 中国科学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究提出Mapping the Concept Landscape框架,以样本级图刻画语义概念全局分布,开发贪心概念覆盖最大化算法,实现更优数据剪枝效率并提供可解释的审计轨迹。
AI 中文摘要
现有数据剪枝方法主要依赖高维特征嵌入来衡量样本重要性,但这些压缩向量往往会掩盖细粒度语义交互,导致剪枝子集对稀有语义概念的覆盖不足。本文提出Mapping the Concept Landscape(MCL,概念图谱绘制),一种用于透明数据剪枝的新型结构感知框架。我们不采用抽象嵌入,而是将每个图像-文本对表示为包含实体、事件和属性的显式样本级图,再将这些独立图整合为全面的数据集级图,以此刻画语义概念的全局分布并量化其在整个语料库中的稀有度。基于这种结构化感知,我们开发了一种贪心概念覆盖最大化算法,该算法会迭代选择样本以最大化高价值、代表性不足概念的边际增益。在多个基准上的实验结果表明,我们的方法不仅比现有最优方法实现了更优的剪枝效率,还为选择过程提供了透明且可解释的审计轨迹。
英文摘要
Existing data pruning methods predominantly rely on high-dimensional feature embeddings to measure sample importance. However, these compressed vectors often obscure fine-grained semantic interactions, leading to suboptimal coverage of rare semantic concepts in the pruned subsets. In this paper, we propose Mapping the Concept Landscape (MCL), a novel structural perception framework for transparent data pruning. Instead of abstract embeddings, we represent each image-caption pair as an explicit sample-level graph comprising entities, events, and attributes. By integrating these individual graphs into a comprehensive dataset-level graph, we characterize the global distribution of semantic concepts and quantify their rarity across the entire corpus. Based on this structured perception, we develop a greedy concept-coverage maximization algorithm that iteratively selects samples to maximize the marginal gain of high-value, under-represented concepts. Experimental results on various benchmarks demonstrate that our method not only achieves superior pruning efficiency compared to state-of-the-art methods but also provides a transparent and interpretable audit trail for the selection process.
CommentsECCV 2026