发表机构
Université Laval(拉瓦尔大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ReLaG是一个模态无关的可扩展框架,通过层次潜变量模型和社区检测推断样本关系组,生成独立训练-测试划分,在分子和蛋白质数据上匹配关系感知方法且扩展性更优,并提供无标签分辨率适应和有效规模估计。
AI 中文摘要
当数据集包含相关样本时(如生化研究中常见的情况),随机划分可能产生非独立的训练-测试子集,这会导致过于乐观的泛化估计。在此,我们引入ReLaG,一个模态无关的框架,通过层次潜变量过程对样本相关性进行建模,并利用邻近图和社区检测推断相关样本的组,以生成独立的训练-测试子集。在分子和蛋白质数据集上,ReLaG与现有关系感知方法性能相当,同时扩展性显著更好,能够在以前不切实际的数据集规模下进行划分。我们进一步引入一种无标签程序,将划分分辨率适应于生产数据,使评估与预期的部署环境保持一致。ReLaG推断的组提供了有效数据集规模的廉价估计,支持多样性感知的数据集扩展。ReLaG是开源的,可通过pip install relag安装。
英文摘要
Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets. Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes. We further introduce a label-free procedure that adapts the splitting resolution to production data, aligning evaluation with the intended deployment setting. ReLaG's inferred groups provide a cheap estimate of effective dataset size, enabling diversity-aware dataset scaling. ReLaG is open source and can be installed with pip install relag.