即时权重生成:ARC-1D上的超网络概念验证
On-the-fly Weight Generation: A Hypernetwork Proof of Concept on ARC-1D
浏览论文内容
中文总结 AI 辅助
本研究以ARC-1D为测试平台,证明超网络能从少量演示即时生成专用模型参数,形成结构化权重空间,支持组合泛化及对未见变换的泛化,移除任务标识符可提升泛化。
中文摘要 AI 辅助
通用模型可以从上下文中适应多种任务,而专用模型可以用更少的容量执行单个功能。然而,获得这样的专用模型需要针对特定任务的训练或适应。我们提出疑问:能否直接从少量演示中生成这些专用模型?以ARC-1D作为受控测试平台,我们展示了单个变换可以由微小的专用模型表示,并且超网络可以从上下文中生成它们的参数。生成的参数形成结构化的权重空间,而由此产生的专用模型表现出部分组合泛化能力,以及对训练中未见过的变换的泛化能力。在这两种设置中,移除显式任务标识符都能改善超越训练变换的泛化能力。综合来看,这些结果提供了一个概念验证:少样本任务上下文可以被即时编译成紧凑的可执行模型参数,并且由此产生的权重空间可以支持已知函数之外的复用和泛化。
英文摘要
General-purpose models can adapt to many tasks from context, while specialised models can execute individual functions with less capacity. Yet obtaining such specialists requires task-specific training or adaptation. We ask whether they can instead be generated directly from a few demonstrations. Using ARC-1D as a controlled testbed, we show that individual transformations can be represented by tiny specialist models, and that a hypernetwork can generate their parameters from context. The generated parameters form a structured weight space, while the resulting specialists show partial compositional generalisation and generalisation to transformations not seen during training. In both settings, removing explicit task identifiers improves generalisation beyond the training transformations. Together, these results provide a proof of concept that few-shot task context can be compiled on-the-fly into compact executable model parameters, and that the resulting weight space can support reuse and generalisation beyond known functions.
发表机构
- University of Oxford(牛津大学)
机构由 AI 辅助整理,请以论文原文为准。