arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.22263cs.CVcs.AI

基于校准残差解码的无训练VLM个性化方法

Training-Free VLM Personalization via Calibrated Residual Decoding

Jiaao Yu, Yujian Ma, Xianming Hu, Pengran Wang, Ang Li

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出无训练校准残差解码框架,通过构建三类证据条件估计个性化边际贡献,经实验在多数据集上提升了VLM的个性化多模态理解性能。

中文摘要 AI 辅助

视觉语言模型(VLM)可在推理时通过直接提供用户画像、偏好或视觉参考实现无训练个性化,无需更新模型参数。但直接的个性化提示无法保证模型可靠利用此类证据:在正用户画像下的预测分布常混合两类来源,一是当前画像真实支持的个性化信号,二是模型的通用视觉或语言先验。因此仅从正画像响应难以判断高置信度答案是由用户画像支持,还是仅反映模型默认偏好。为解决该问题,本文提出一种无训练校准残差解码框架:给定相同图像与问题,构建正画像、反事实画像、空画像三类证据条件;该方法以某一条件下的预测为锚定基准,通过三类条件下的分数差显式估计个性化的边际贡献;还引入基于归一化熵的不确定性校准,使个性化增强强度可适应残差信号的可靠性。在MMPB、YoLLaVA、MyVLM数据集上的实验表明,所提方法无需微调即可提升个性化多模态理解,在身份敏感的视觉个性化任务上取得一致增益;额外分析显示,当对比个性化信号不确定时,熵校准可稳定残差解码过程。

英文摘要

Vision-language models can be personalized in a training-free manner by directly providing user profiles, preferences, or visual references at inference time, without updating model parameters. However, direct personalized prompting does not guarantee that the model will reliably exploit such evidence. The predictive distribution under the positive user profile often mixes two sources: personalized signals genuinely supported by the current profile, and the model's generic visual or linguistic priors. As a result, from the positive-profile response alone, it is difficult to determine whether a high-confidence answer is supported by the user profile or merely reflects the model's default preference. To address this problem, we propose a training-free calibrated residual decoding framework. Given the same image and question, we construct three evidence conditions: a positive profile , a counterfactual profile , and an empty profile . Our method keeps the prediction under as the anchored base, and explicitly estimates the marginal contribution of personalization from score differences across the three conditions. We further introduce normalized-entropy-based uncertainty calibration, allowing the strength of personalized enhancement to adapt to the reliability of the residual signal. Experiments on MMPB, YoLLaVA, and MyVLM show that the proposed method improves personalized multimodal understanding without fine-tuning, with consistent gains on identity-sensitive visual personalization tasks. Additional analysis shows that entropy calibration stabilizes residual decoding when the contrastive personalization signal is uncertain.

↑