arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.17004cs.CV

对称感知的似然轨道聚合用于选择性左右声明验证

Symmetry-Aware Likelihood-Orbit Aggregation for Selective Left-Right Claim Verification

发表机构浙江大学工程师学院 · 浙江大学电气工程学院 · 浙江大学海洋学院
另 1 家 · 查看机构详情
  • Polytechnic Institute, Zhejiang University(浙江大学工程师学院)
  • College of Electrical Engineering, Zhejiang University(浙江大学电气工程学院)
  • Ocean College, Zhejiang University(浙江大学海洋学院)
  • College of Artificial Intelligence, Zhejiang University(浙江大学人工智能学院)

机构由 AI 辅助整理,请以论文原文为准。

Zhouzhi Xiong, Chuxi Zhang, Weizhen He, Yi Chen, Qi Li, Donglian Qi

首次发表
浏览论文内容

中文总结 AI 辅助

针对冻结视觉语言模型在左右声明验证上的不可靠性,提出Relation-Orbit闭式对比方法,通过八种归一化似然度组合实现选择性验证,在VSR和GQA上优于基线,提升覆盖率。

中文摘要 AI 辅助

冻结的视觉语言模型(VLM)在细粒度的左右声明上仍然不可靠,而原始声明似然度并不一定能可靠地对验证错误进行排序。在固定水平反射干预后,如何将其引发的似然测量组合成选择性验证信号?我们提出Relation-Orbit,一种无学习融合参数的闭式对比方法,将八个归一化似然度分配给由反射、逆关系和实体交换确定的查询支持和反事实角色。仅当符号对比度超过基于保留数据使用逐点Clopper-Pearson上置信界选择的阈值时,才断言声明。在VSR和GQA上跨四个冻结VLM,Relation-Orbit在10%选择性风险校准目标下,在所有八个数据集-骨干设置中均比全八轨道的Orbit-Max基线获得更高的平均测试覆盖率;对几乎全部弃权(不执行)的单侧干预分数的增益单独报告。单独的LLaVA-1.5/COCO评估、缩减轨道对照和双侧分区诊断进一步表征了结构优势。

英文摘要

Frozen vision-language models (VLMs) remain unreliable on fine-grained left-right claims, and raw claim likelihoods need not reliably rank verification errors. After a horizontal-reflection intervention is fixed, how should its induced likelihood measurements be combined into a selective verification signal? We introduce Relation-Orbit, a closed-form contrast with no learned fusion parameters that assigns eight normalized likelihoods to query-supporting and counterfactual roles determined by reflection, inverse relation, and entity exchange. A claim is asserted only when the signed contrast exceeds a threshold selected on held-out data using pointwise Clopper-Pearson upper confidence bounds. On VSR and GQA across four frozen VLMs, Relation-Orbit yields higher mean test coverage at a 10% selective-risk calibration target than an all-eight Orbit-Max baseline in all eight dataset-backbone settings; gains over a nearly abstain-all one-sided intervention score are reported separately. A separate LLaVA-1.5/COCO evaluation, reduced-orbit controls, and a two-sided partition diagnostic further characterize the structural advantage.

补充信息

↑