arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

一种用于稀疏计数矩阵的具有零训练特征提取的确定性二进制指纹框架

A Deterministic Binary Fingerprinting Framework with Zero-Trained Feature Extraction for Sparse Count Matrices

Lei Zhao, Fujin Huang, Ling Kang, Quan Guo

arXiv 2607.14596首次发表:更新:

AI 中文总结

研究针对稀疏计数矩阵,提出MMTB确定性二进制指纹框架,无需标签监督等。经列归一化和固定阈值映射样本,其汉明距离与L1距离有对应关系。在实验中,该框架在特定任务表现良好,适用于粗粒度分离与资源受限部署,还能减少内存。

AI 中文摘要

稀疏计数矩阵,从单细胞转录组到k-mer图谱和文档-词频,传统上通过主成分分析(PCA)降维的图聚类或连续嵌入空间中的迭代优化进行分析。我们引入了MMTB,这是一个确定性的无学习二进制表示框架,不需要标签监督、模型拟合或基于梯度的优化。按列进行最小-最大归一化,然后通过固定阈值将每个样本映射到一个温度计指纹,其汉明距离与归一化的L1距离显示出经验对应关系,在单细胞RNA-seq对上皮尔逊相关系数约为0.92。在一个有利的三细胞系混合物中,三阈值指纹在每个细胞188字节时实现了0.99的互信息(NMI)。在公平的汉明最近邻图加上莱顿读出的情况下,MMTB在这个粗略任务上接近PCA加上莱顿。在具有挑战性的组织样注释上,连续管道通常表现更好;对于PBMC Seurat,MMTB的NMI为0.39,而Scanpy为0.49,这突出表明MMTB适用于粗粒度分离和资源受限的部署,而不是细粒度亚型发现或作为连续嵌入的一般替代品。相对于密集的float32表示,MMTB指纹将内存减少了约10倍,同时提供固定宽度的可汉明索引的代码。提供了一个无标签适用性分数作为部署指南,而不是性能预测器。

英文摘要

Sparse count matrices from single-cell transcriptomes to k-mer profiles and document-term frequencies are conventionally analyzed via PCA-reduced graph clustering or iterative optimization in continuous embedding spaces. We introduce MMTB, a deterministic non-learned binary representation framework that requires no label supervision, model fitting, or gradient-based optimization. Column-wise Min-Max normalization followed by fixed cutoffs maps each sample to a thermometer fingerprint whose Hamming distances show empirical correspondence with normalized L1 distances, with Pearson correlation approximately 0.92 on single-cell RNA-seq pairs. In a favorable three-cell-line mixture, the 3-threshold fingerprint achieves NMI of 0.99 at 188 bytes per cell. Under a fair Hamming nearest-neighbor graph plus Leiden readout, MMTB approaches PCA plus Leiden on this coarse task. On challenging tissue-like annotations, continuous pipelines often lead; PBMC Seurat NMI is 0.39 for MMTB versus 0.49 for Scanpy, underscoring that MMTB is suited for coarse-grained separation and resource-constrained deployments rather than fine-grained subtype discovery or as a general replacement for continuous embeddings. Relative to dense float32 representations, MMTB fingerprints reduce memory by approximately 10-fold while providing fixed-width Hamming-indexable codes. PCA30 embeddings and sparse CSR may be smaller; we do not claim universal compression. A label-free suitability score is provided as a deployment guideline, not a performance predictor.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑