AdaKerNet:面向多模态大模型任务自适应预测的神经核解码
AdaKerNet: Neural Kernel Decoding for Task-Adaptive Prediction with Multimodal Large Models
查看机构详情
- University of California San Diego(加州大学圣迭戈分校)
机构由 AI 辅助整理,请以论文原文为准。
浏览论文内容
中文总结 AI 辅助
AdaKerNet提出一种不依赖MLLM参数的可学习神经核解码器,通过Lipschitz控制特征、参考核与轻量神经预测器联合优化,在稀缺标签下跨四种多模态模型平均错误率降低达41%。
中文摘要 AI 辅助
大型基础模型以其高效适应下游任务的承诺被引入。然而,在有限监督下,作为大型基础模型重要一类的多模态大语言模型(MLLMs)仍难以适应各种下游任务。适应通常依赖于MLLM参数微调或训练基于神经的解码器。这两种方法在有限监督下都面临困难,而微调还需要访问模型参数,这对于闭源模型通常不可用。我们引入了AdaKerNet,一种新颖的可学习的任务自适应神经核解码器。AdaKerNet完全不依赖于底层MLLM的参数,仅基于从多种可用模态获得的(冻结的)丰富表示进行操作。AdaKerNet依赖于(i)一组从这些MLLM表示中导出的可学习的、Lipschitz控制的多模态特征;(ii)一个参考核,为这些特征提供软结构先验;以及(iii)一个轻量级非线性神经预测器,自适应地变形该结构。在统一优化框架内联合学习核表示和神经预测器,使AdaKerNet能够捕获与下游任务相关的特征和几何关系。在四个MLLM(BLIP-2、LLaVA-1.5、Qwen2.5-VL和Gemini Embedding 2)以及涵盖文本、音频、图像和表格测量的多模态输入上的数值测试表明,在多种稀缺标签预算下,与直接MLP、基于注意力、自编码器和基于核的解码器相比,AdaKerNet实现了显著且一致的改进,跨基线的平均错误率降低高达41%。这些结果确立了AdaKerNet在稀缺标签机制下从冻结多模态表示进行预测的有效方法。额外的结构消融研究突出了AdaKerNet各组件的互补贡献。
英文摘要
Large foundation models have been introduced with the promise of efficient adaptation to downstream tasks. Yet, under limited supervision, MLLMs, an important class of large foundation models, remain challenging to adapt to various downstream tasks. Adaptation typically relies either on MLLM parameter fine-tuning or on training neural-based decoders. Both approaches struggle under limited supervision, while fine-tuning additionally requires access to model parameters, which is often unavailable for closed-source models. We introduce AdaKerNet, a novel learnable task-adaptive neural kernel decoder. AdaKerNet is fully agnostic to the parameters of the underlying MLLM and operates solely on its (frozen) rich representations obtained from the diverse available modalities. AdaKerNet relies on (i) a set of learnable, Lipschitz-controlled multimodal features derived from these MLLM representations; (ii) a reference kernel that provides a soft structural prior on those features; and (iii) a lightweight nonlinear neural predictor that adaptively deforms that structure. Learning the kernel representation and the neural predictor jointly within a unified optimization framework allows AdaKerNet to capture features and geometric relationships relevant to the downstream task. Numerical tests across four MLLMs: BLIP-2, LLaVA-1.5, Qwen2.5-VL, and Gemini Embedding 2, and multimodal inputs spanning text, audio, images, and tabular measurements demonstrate significant and consistent improvements over direct MLP, attention-, autoencoder- and kernel-based decoders, across a range of scarce-label budgets, with average error reduction of up to 41% across baselines. These results establish AdaKerNet as an effective approach for prediction from frozen multimodal representations in the scarce label regime. Additional structural ablations highlight the complementary contributions of AdaKerNet's components.