arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

几何在发挥作用吗?对双曲视觉-语言模型层次性的工作点审计

Is the Geometry Doing the Work? An Operating-Point Audit of Hierarchy in Hyperbolic Vision-Language Models

Jaeyoung Kim, Eunseok Kim, Dongsuk Jang

arXiv 2607.05268首次发表:更新:

发表机构

MADI(MADI)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对双曲视觉-语言模型是否利用其几何特性的问题,提出多组诊断指标审计三类主流模型,发现其实际未激活双曲径向/锥机制,层次性表现与几何特性无关。

AI 中文摘要

双曲表示模型是否利用其自身几何特性无法仅通过曲率参数判断:关键在于无量纲工作点$\sqrt{c}\rho$,以及径向和锥机制是否在该点激活。我们开发了一系列必要条件诊断工具,对三个已发表的双曲视觉-语言模型系列——MERU、HyCoCLIP和PHyCLIP——在公开检查点和固定GRIT快照上的受控干预下开展审计,识别出三类失效模式:第一,曲率并非有效可用资源:工作点始终接近欧氏区域($H(u)\approx 1$;所有经审计的收敛检查点均未达到$\sqrt{c}\rho>1$),解除曲率下限约束会改变曲率和范数,但工作点仍保持近欧氏状态,下游性能无明显下降。第二,锥和遍历机制经检测未生效:蕴含锥处于未激活、饱和或错位状态,分级遍历在受控读出条件下失效,定向径向深度在量化灵敏度下仅为高于打乱空对照的有界非检测值,唯一留存的原生关系残差也不具备可操作性。第三,看似层次性的评估结果不具备决定性:分类学相关性由角距离承载,粗检索性能提升与框/组合监督相关,而非曲率。我们给出了机理解释:蕴含目标允许存在低曲率、宽锥的捷径,且一个无参数孔径恒等式(锥饱和当且仅当$\sqrt{c}\rho\le 2K$)定位了所有经蕴含训练的无约束运行的收敛边界;未采用蕴含训练的运行不会在此处停滞。该捷径是模型坍缩的主要加速因素,但并非唯一原因。这些已发布的模型实现并未体现其几何设计所预期的径向/锥机制;我们将审计方法提炼为面向未来层次性相关声明的五维几何报告框架。

英文摘要

Hyperbolic vision-language models are designed to encode abstraction geometrically: general concepts near the origin, specific ones farther out, and entailment cones representing directed order. We ask whether trained MERU, HyCoCLIP, and PHyCLIP models actually use these mechanisms. We audit seven released checkpoints and matched from-scratch interventions, using diagnostics that distinguish active hyperbolic geometry from angular structure and supervision effects. All audited converged checkpoints remain near-Euclidean in the dimensionless radius $u=\sqrt{c}ρ$, which measures how strongly embeddings experience hyperbolic geometry: the largest observed image-side value is $0.37$ -- well below $u\approx0.84$, where local metric distortion reaches $10\%$. Releasing the curvature floor changes curvature and norms but not this regime, with mixed, generally modest downstream shifts. Trained entailment cones are saturated or nearly saturated, so low violation rates can arise from trivially wide cones rather than learned order. Preregistered semantic traversal detects weak within-branch order but no operative full-hierarchy readout. Shuffle-controlled tests detect no pair-specific radial ordering in released checkpoints, and no positive result is consistent across all three matched ViT-B seeds. We trace this to a low-curvature shortcut: lowering curvature widens entailment cones and suppresses violations without learning order. In the probed trajectories, gradient decomposition identifies entailment as the dominant curvature-lowering pressure during collapse. Yet curvature contracts even when entailment is removed, so the shortcut is not the sole cause. Under our diagnostics, the audited formulations do not demonstrate an operative radial or cone-based hierarchy. We distill the audit into a five-number geometry report for evaluating future hierarchy claims.

Comments48 pages, 5 figures, Under review at TMLR

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑