复杂度诱导:通过结构化标签失真实现组合泛化
Complexity Induction: Compositional Generalization via Structured Training Distortion
AI总结:
研究提出复杂度诱导方法,通过结构化标签失真(混合标签、扩展数据集),使标准CNN分类器在不修改架构时具备组合泛化能力,可预测未见颜色-形状组合,为自然语言在认知发展中的作用提供启发。
AI中文摘要:
我们证明,对训练数据进行结构化失真(我们称之为复杂度诱导),可在不修改架构的标准CNN分类器中诱导出组合泛化能力。使用彩色几何形状的合成图像,我们将类别编码为无显式属性分解的扁平字符串标签(例如“red-circle”),并完全排除某些颜色-形状组合用于训练。我们采用两种基于类别名称间Jaccard字符串相似度的失真方法:混合标签(编码类间重叠的软目标分布)和扩展数据集(带有结构动机错误标签的虚假训练样本)。两种方法均能诱导对未见类别组合的预测能力,且作用于不同层面:混合标签通过利用CNN的自然嵌入结构激活分类器对未见组合的处理,而扩展训练则改善嵌入分解本身。采用随机(非结构化)错误标签的对照实验证实,该效果取决于失真的结构,而非单纯的噪声。这些结果表明,训练信号的结构化复杂化可同时影响学习表征的内部组织及其组合解释——这一原理可能是自然语言在认知发展中作用的基础。
英文摘要:
We demonstrate that structured distortion of training data - which we term complexity induction - can induce compositional generalization in a standard CNN classifier without architectural modification. Using synthetic images of colored geometric shapes, we encode classes as flat string labels (e.g., "red-circle") with no explicit attribute decomposition, and exclude certain color-shape combinations from training entirely. We apply two distortion methods derived from Jaccard string similarity between class names: mixed labels (soft target distributions encoding inter-class overlap) and expanded dataset (false training samples with structurally motivated incorrect labels). Both methods induce the ability to predict unseen class combinations, and act at different levels: mixed labels activate the classifier for unseen combinations by exploiting the CNN's natural embedding structure, while expanded training improves the embedding factorization itself. A control with random (unstructured) false labels confirms that the effect depends on the structure of the distortion, not on noise per se. These results suggest that structured complication of training signals can influence both the internal organization of learned representations and their compositional interpretation - a principle that may underlie the role of natural language in cognitive development.