发表机构
University of Minnesota; IBM Research AI(明尼苏达大学; IBM研究院人工智能)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出分布性超网络,根据查询预测LoRA权重更新分布,通过蒙特卡洛采样实现测试时扩展,性能优于确定性超网络及token采样基线。
AI 中文摘要
超网络最近在根据任务描述或额外演示等信号,在运行时动态调整大型语言模型(LLM)的参数方面取得了成功。在此,我们提出疑问:仅使用输入给LLM的查询,能获得多少自适应信号?为回答此问题,我们研究了用于LoRA估计的查询条件超网络。此外,我们引入了分布性超网络,不仅能产生参数适配器的点估计,还能产生可能的LoRA分布。为此,我们提出了一种使用可微分蒙特卡洛近似的简单端到端损失,并探索了多种分布参数化方法,包括回归和凸组合变体。结果表明,即使使用学习分布的平均值也能优于确定性超网络。关键的是,学习到的分布实现了一种不同形式的测试时扩展:不是仅通过从固定模型采样更多token序列来增加额外计算,而是采样权重更新,为同一查询生成多个自适应模型。随着考虑的权重样本增多,性能得到提升,并且仍强于相应的token采样自适应基线。最后,我们发现生成的更新可以在查询间转移,这表明超网络学习了模型应如何自适应的可重用结构。总之,这些结果表明,查询条件下的权重更新分布既能支持自适应,也能支持测试时扩展。
英文摘要
Hypernetworks have recently shown success in dynamically adapting the parameters of Large Language Models (LLMs) at runtime based on signals such as task descriptions or additional demostrations. Here we ask: how much adaptation signal can be obtained using only the input query to an LLM?. To answer this, we study query-conditioned Hypernetworks for LoRA estimation. Further, we introduce distributional Hypernetworks, able to produce not only point estimates of parameter adaptors, but also a distribution over possible LoRAs. For this we propose a simple end-to-end loss using a differentiable Monte Carlo approximation and explore multiple distribution parametrizations including regression and convex combination variants. Results show that even using the mean of the learned distribution can outperform deterministic hypernetworks. Crucially, the learned distribution enables a different form of test-time scaling: instead of spending additional compute only by sampling more token sequences from a fixed model, we sample weight updates, yielding multiple adapted models for the same query. Performance improves as more weight samples are considered and remains stronger than corresponding token-sampling adaptation baselines. Finally, we find that generated updates can transfer across queries, suggesting that the hypernetwork learns reusable structure in how the model should adapt. Together, these results show that query-conditioned distributions over weight updates can support both adaptation and test-time scaling.