arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

不同分解与含异常随机数据集欠采样的临界性

Criticality in Dissimilar Decomposition and Undersampling of Random Datasets with Anomalies

Ghurumuruhan Ganesan

arXiv 2609.13201首次发表:更新:

发表机构

University of Bristol(布里斯托大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本文研究含异常随机数据集的强不同分解与欠采样临界性,证明相变现象并给出规模界限,为LLM训练数据分解提供理论指导。

AI 中文摘要

未来大型语言模型的训练数据集将包含大量由当前LLM生成的AI文本/图像数据。在这种情况下,理解这如何影响批次分解以及由此产生的新LLM的性能至关重要。本文将AI生成的数据视为与“主”数据点“关联”的异常,并研究整体随机数据集的分解和欠采样性质。我们使用冗余图和迭代技术来获得强不同(SD)分解最小规模的界限,并证明了一个相变现象:当异常数量较少时,最小规模主要由主数据点决定,而当异常超过某个阈值时,该最小规模被异常“接管”。我们还建立了随机欠采样数据集强相似性的规模临界性结果,并通过涉及类别数据集的示例来说明我们的结果,这些数据集的空间规模远大于数据集本身的规模。

英文摘要

Training datasets for upcoming LLMs would include a significant amount of AI text/image data generated from current LLMs. In such a scenario, it is important to understand how this affects batch decompositions and thereby, the performance of the resultant new LLM. In this paper, we consider AI generated data as anomalies ``linked" to main data points and study decomposition and undersampling properties of the overall random dataset. We use redundancy graphs and iteration techniques to obtain bounds for the minimum size of a strongly dissimilar (SD) decomposition and demonstrate a phase transition phenomena, wherein the minimum size is essentially determined by the \emph{main} data points when the number of anomalies is small and is ``taken" over by the anomalies above a certain threshold. We also establish a size criticality result for the strong similarity of a randomly undersampled dataset and illustrate our results with examples involving categorical datasets, whose overall space size is much larger than the size of the dataset.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑