arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17040cs.AI

稀疏MLLM锚点,密集适配:打破野外测试时适配中的自引用循环

Sparse MLLM Anchors, Dense Adaptation: Breaking the Self-Referential Loop in Wild Test-Time Adaptation

Zhenbin Wang, Lei Zhang, Lituan Wang, Yan Wang, Zhao Zhang, Wei Huang

首次发表
浏览论文内容

中文总结 AI 辅助

针对野外测试时适配中模型自引用循环问题,提出MASA方法,利用冻结多模态大语言模型为少量锚点生成语义描述,并传播至邻近样本,通过在线原型记忆辅助轻量适配,在ImageNet-C基准上验证有效性。

中文摘要 AI 辅助

野外测试时适配(WTTA)在小测试批次、并发分布偏移和时间变化的类别不平衡条件下,在线更新源模型。大多数WTTA方法从被适配的模型中提取适配信号,包括预测不确定性、样本可靠性和局部特征几何。当源模型在偏移下不可靠时,这些信号可能强化其自身错误,形成自引用循环。我们提出MASA(多模态大语言模型锚定的语义适配),该方法用冻结的多模态大语言模型(MLLM)的结构化语义描述补充模型内部证据。为限制推理成本,MASA仅对一小部分多样化、按可靠性排序的锚点查询MLLM。生成的描述捕获对象类别以及风格、视角和遮挡等干扰因素。MASA对这些描述进行编码,将其传播到邻近测试样本,并将得到的视觉-语义信息存储在在线原型记忆中。从该记忆进行描述符感知检索为轻量适配归一化仿射参数提供辅助目标。我们在WTTA ImageNet-C基准上,使用ResNet和ViT骨干网络,在有限批次、混合域和类别不平衡标签偏移设置下评估MASA。

英文摘要

Wild test-time adaptation (WTTA) updates a source model online under small test batches, concurrent distribution shifts, and time-varying class imbalance. Most WTTA methods derive their adaptation signals, including predictive uncertainty, sample reliability, and local feature geometry, from the model being adapted. When the source model is unreliable under shift, these signals can reinforce its own errors, forming a self-referential loop. We introduce MASA (Multimodal-LLM-Anchored Semantic Adaptation), which complements model-internal evidence with structured semantic descriptions from a frozen multimodal large language model (MLLM). To limit inference cost, MASA queries the MLLM only for a small set of diverse, reliability-ranked anchors. The resulting descriptions capture the object family and nuisance factors such as style, viewpoint, and occlusion. MASA encodes these descriptions, propagates them to neighboring test samples, and stores the resulting visual-semantic information in an online prototype memory. Descriptor-aware retrieval from this memory provides an auxiliary target for lightweight adaptation of normalization-affine parameters. We evaluate MASA on the WTTA ImageNet-C benchmark under limited-batch, mixed-domain, and imbalanced-label-shift settings with ResNet and ViT backbones.

发表机构

  • Sichuan University(四川大学)

机构由 AI 辅助整理,请以论文原文为准。

↑