arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

ReLaG: 一种将随机划分泛化到具有潜在关系数据的可扩展框架

ReLaG: A Scalable Framework Generalizing Random Splits to Data with Latent Relations

Anthony Lavertu, Jacob Cote, Sophie Gobeil, Jacques Corbeil, Isabeau Premont-Schwarz, Pascal Germain

arXiv 2609.38538首次发表:更新:

发表机构

Université Laval(拉瓦尔大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

ReLaG是一个模态无关的可扩展框架,通过层次潜变量模型和社区检测推断样本关系组,生成独立训练-测试划分,在分子和蛋白质数据上匹配关系感知方法且扩展性更优,并提供无标签分辨率适应和有效规模估计。

AI 中文摘要

当数据集包含相关样本时(如生化研究中常见的情况),随机划分可能产生非独立的训练-测试子集,这会导致过于乐观的泛化估计。在此,我们引入ReLaG,一个模态无关的框架,通过层次潜变量过程对样本相关性进行建模,并利用邻近图和社区检测推断相关样本的组,以生成独立的训练-测试子集。在分子和蛋白质数据集上,ReLaG与现有关系感知方法性能相当,同时扩展性显著更好,能够在以前不切实际的数据集规模下进行划分。我们进一步引入一种无标签程序,将划分分辨率适应于生产数据,使评估与预期的部署环境保持一致。ReLaG推断的组提供了有效数据集规模的廉价估计,支持多样性感知的数据集扩展。ReLaG是开源的,可通过pip install relag安装。

英文摘要

Random splitting can yield non-independent train--test subsets when a dataset contains related samples, as is common in certain applications such as biochemical studies. This leads to overly optimistic generalization estimates. Here, we introduce ReLaG, a modality-agnostic framework that models sample relatedness through a hierarchical latent-variable process and infers groups of related samples using proximity graphs and community detection to produce independent train--test subsets. Across molecular and protein datasets, ReLaG matches existing relation-aware methods while scaling substantially better, enabling splits at previously impractical dataset sizes. We further introduce a label-free procedure that adapts the splitting resolution to production data, aligning evaluation with the intended deployment setting. ReLaG's inferred groups provide a cheap estimate of effective dataset size, enabling diversity-aware dataset scaling. ReLaG is open source and can be installed with pip install relag.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑