arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

PAIQ:通过残差旋转进行补丁对齐的语义注入

PAIQ: Patch-Aligned Semantic Injection via Residual Rotation

Pinze Ren, Yuwei Zhang, Hao Chen, Linghao Meng, Chang Li, Qiankun Li

arXiv 2609.37685首次发表:更新:

发表机构

Tsinghua University; Beijing University of Posts and Telecommunications; National University of Singapore; Nanyang Technological University(清华大学; 北京邮电大学; 新加坡国立大学; 南洋理工大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

PAIQ通过正交旋转注入补丁对齐语义,融合DINOv3和SigLIP特征,提升多模态理解正确性并减少幻觉。

AI 中文摘要

语言对齐和自监督的视觉编码器在语义抽象和空间细节方面提供了互补的优势。利用这种互补性需要在丰富局部特征的同时,保留语义相关补丁之间的区分度。我们提出了PAIQ,一种补丁对齐的语义注入框架,它结合了基于内容的跨编码器匹配与正交约束的残差更新。使用DINOv3补丁特征作为空间基础,PAIQ通过联合源分配聚合互补的SigLIP特征,并通过共享的正交变换Q注入聚合与基础的差异。这种旋转适应更新方向,同时保留残差范数和成对角度。对于固定的投影特征,我们推导了在相似语义聚合下补丁可分离性的条件,并表明当聚合共享时,旋转在直接插值上增加了非负的分离项。仅训练投影和融合参数;两个视觉编码器和语言模型保持冻结,融合保留196个视觉标记。在多样化的语言骨干网络上,与单编码器接口相比,PAIQ在图像描述和视觉问答上,在评判者评估的正确性方面取得了广泛的提升,并减少了幻觉严重程度。在2B和9B的Qwen骨干网络上,这种紧凑接口平均比最强的评估融合或标记压缩基线高出约2.9个正确性点。

英文摘要

Language-aligned and self-supervised visual encoders offer complementary strengths in semantic abstraction and spatial detail. Harnessing this complementarity requires enriching local features while retaining distinctions between semantically related patches. We introduce PAIQ, a patch-aligned semantic injection framework that combines content-based cross-encoder matching with orthogonally constrained residual updates. Using DINOv3 patch features as the spatial base, PAIQ aggregates complementary SigLIP features through joint source allocation and injects the aggregate--base differences through a shared orthogonal transformation Q. This rotation adapts update directions while preserving residual norms and pairwise angles. For fixed projected features, we derive conditions for patch separability under similar semantic aggregates and show that rotation adds a nonnegative separation term over direct interpolation when the aggregate is shared. Only the projection and fusion parameters are trained; both visual encoders and the language model remain frozen, and fusion retains 196 visual tokens. Across diverse language backbones, PAIQ yields broad gains in judge-assessed correctness and reductions in hallucination severity over single-encoder interfaces on image description and visual question answering. On the 2B and 9B Qwen backbones, this compact interface outperforms the strongest evaluated fusion or token-compression baselines by about 2.9 correctness points on average.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑