使单细胞数据提炼可审计:通过离散最小-最大选择实现可追溯的真实细胞核心集
Making Single-Cell Data Distillation Auditable: Traceable Real-Cell Coresets via Discrete Min--Max Selection
- The Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳))
- Shenzhen Loop Area Institute(深圳河套学院)
- Shenzhen Research Institute of Big Data(深圳大数据研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究单细胞数据存储、审计和复用成本高的问题,提出固定CF和Minmax-CF两种真实细胞选择器,通过离散最小-最大选择实现可追溯细胞核心集,在多数据集实验中表现良好,能减少最差方向差异,不同数据集和任务下下游效用和成本有别。
AI中文摘要:
单细胞数据集在存储、审计和用于模型训练的复用方面成本日益增加。降维和数据集提炼可减轻负担,但传统提炼方法常产生无法追溯到被检测细胞的合成表达谱。我们将可追溯的单细胞数据提炼定义为在固定的细胞和基因预算下保留原始细胞标识符和基因符号。由此产生的训练子集仍与测量的计数、标签和检测元数据相关,可据此检查意外预测。我们提出了两种真实细胞选择器。固定CF使用静态特征函数匹配。Minmax-CF解决一个熵正则化的离散最小-最大问题,该问题对保存不佳的方向加权,并仅添加观察到的细胞。在三个数据集上的供体、技术和扰动水平变化中,Minmax-CF在MS上保留了96.52%的全平衡准确率,在hPancreas上平均与全数据集近似匹配,在全基因设置下GPU加速中位数为2.55倍,在Norman上压缩方法中获得最低通路误差。对于罕见状态、一些技术变化、未见过的扰动成分以及保真度与下游效用弱相关的设置,性能仍然较弱。由于所选ID指的是被检测细胞,这些情况可通过检查相应的训练支持、标签和检测元数据来研究。Minmax-CF持续减少最差方向差异,而下游效用和成本因数据集和任务而异。
英文摘要:
Large single-cell datasets are expensive to store, curate, and repeatedly reuse for model training. Data distillation can reduce this burden by building smaller training sets. However, many existing methods rely on synthetic cells. These synthetic cells do not retain direct correspondence with assayed cells and genes. This limits source-level inspection and biological traceability. Moreover, real-cell expression matrices are often sparse and noisy. In light of these challenges, we propose Minmax-CF, a label-aware characteristic-function selector for traceable single-cell data distillation. Minmax-CF formulates compression as a discrete min--max selection problem over characteristic-function directions. It uses entropy-regularized maximization to emphasize the least preserved directions. Greedy minimization ranks cells and genes by how much they reduce the resulting weighted error. The method alternates cell and gene selection under explicit axis-specific budgets. Across five coarse-lineage benchmarks and five compression budgets, Minmax-CF retains 95.3% of the Full-reference macro-F1 on average, with gaps that exceed one per-seed standard deviation. It also retains exact source-cell indices and original gene symbols. Compared with size-matched synthetic PCA-Centroid and Distribution Matching (DM) baselines, Minmax-CF achieves higher coarse-lineage macro-F1 in 24 of 25 comparisons against each baseline. It exceeds their average performance by 10.4% and 17.4%, respectively. Retained cells can also be projected onto independently computed embeddings for direct biological interpretation.