arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当提示忽略结构时:基于图的属性推理用于校准视觉语言模型

When Prompts Ignore Structure: Graph-Based Attribute Reasoning for Calibrated VLMs

Tanay Sodha, Aditya Sharma, Ramya Hebbalaguppe, Vinti Agarwal, Pranav Murthy Yeluripaty

arXiv 2607.07395首次发表:更新:

发表机构

Birla Institute of Technology and Science, Pilani; TCS Research(贝拉科技与科学学院皮拉尼分校; 塔塔咨询服务公司研究院)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究视觉语言模型测试时自适应中可靠置信度估计问题,提出ARGTCA方法,将(类,属性)对表示为符号属性图节点,用对比目标训练图注意力网络,引入两种属性选择策略,实验证明该方法能有效降低校准误差。

AI 中文摘要

可靠的置信度估计仍然是视觉语言模型(VLM)测试时自适应的关键限制,提示调整提高了零样本准确率,但由于熵驱动的过度自信常常会降低校准。先前方法使用基于语言模型的类属性和对比正则化来缓解此问题,但独立处理属性,忽略其关系结构。我们提出ARGTCA,将(类,属性)对表示为符号属性图中的节点,并使用对比目标训练图注意力网络(GAT)以生成捕获属性间依赖关系的结构信息嵌入。我们引入两种属性选择策略:用于类内多样性的ARGTCA-DIV和用于类间区分的ARGTCA-DISC。九个基准测试的实验表明,ARGTCA-DIV比基线平均降低约37%的预期校准误差(ECE),而ARGTCA-DISC始终是第二好的变体,比基线平均降低约17%的ECE。这些结果表明,对符号属性交互进行建模为VLM中可靠的测试时自适应提供了一种原则性方法。

英文摘要

Reliable confidence estimation remains a key limitation of test-time adaptation in vision-language models (VLMs), where prompt tuning improves zero-shot accuracy but often degrades calibration due to entropy-driven overconfidence. Prior approaches mitigate this using LLM-derived class attributes and contrastive regularization, yet treat attributes independently, ignoring their relational structure. We propose ARGTCA, which represents (class, attribute) pairs as nodes in a Symbolic Attribute Graph and trains a Graph Attention Network (GAT) using contrastive objectives to produce structurally informed embeddings that capture inter-attribute dependencies. We introduce two attribute selection strategies: ARGTCA-DIV for intra-class diversity and ARGTCA-DISC for inter-class discrimination. Experiments across nine benchmarks show that ARGTCA-DIV reduces average Expected Calibration Error (ECE) by approximately ~37% over baselines, while ARGTCA-DISC consistently performs as the second-best variant, reducing average ECE by approximately ~17% over baselines. These results suggest that modeling symbolic attribute interactions provides a principled approach for reliable test-time adaptation in VLMs.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑