跨域学习中需要多少样本才足够?
How Many Samples Are Enough for Learning Across Domains?
AI总结:
本研究建立每域样本需求标准,揭示训练域数量与每域样本数间的反比线性规律,为数据充分性提供理论指导,并证明域内学习与域外泛化的紧密关系。
AI中文摘要:
理解学习的基本机制对于设计具有强泛化能力的系统至关重要。近期研究表明,当每个域包含足够多的数据样本时,增加训练域的数量或扩大它们之间的分布偏移,可以改善泛化性能。然而,数据样本在何种条件下可被视为足够,仍未被探索。在本工作中,我们通过基于所提出的学习界建立每域样本需求的标准,填补了这一空白。这些标准不仅揭示了训练域数量与每域所需样本数量之间的反比线性缩放规律,还解释了数据充分性假设背后的基本原理,从而为评估现有数据集的充分性和构建数据集提供了理论指导。这与经典学习理论不同,因为所需样本数量高度依赖于训练域的数量。此外,我们通过所提出的泛化界证明了域内学习与域外泛化之间的密切关系,并最后讨论了一些关键论点。
英文摘要:
Understanding the fundamental mechanisms of learning is essential for designing systems with strong generalization. Recent studies have shown that increasing the number of training domains, or enlarging the distribution shift among them, improves generalization when each domain contains sufficiently many data samples. However, the conditions under which the data samples can be considered sufficient remain unexplored. In this work, we fill this gap by establishing criteria for per-domain sample requirements based on the presented learning bounds. These criteria not only reveal an inverse linear scaling law between the number of training domains and the number of samples required per domain, but also explain the fundamental rationale behind the assumption of data sufficiency, thereby providing theoretical guidance for assessing the adequacy of existing datasets and constructing datasets. This differs from classical learning theory, as the number of samples required is highly dependent on the number of training domains. Additionally, we prove the close relationship between in-domain learning and out-of-domain generalization through the presented generalization bounds, and lastly discuss some key arguments.