发表机构
Zhuhai College of Science and Technology; Zhuhai UNO Technology Co., Ltd.(珠海科技学院; 珠海优诺科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出方向几何响应(DGR),通过位移在体积梯度上的投影解释多模态几何分数对退化的响应,证明方向比幅度更关键,显著优于无方向方法。
AI 中文摘要
基于Gram行列式的几何对齐分数提供了一种紧凑的方式来建模模态间的高阶一致性,然而这些分数对模态退化的响应机制尚不明确。本文探讨多模态几何分数的响应是否主要由扰动引起的位移幅度决定。利用MSR-VTT(N=878)和DiDeMo(N=980)的冻结队列,我们施加受控的视频模糊和音频噪声,并分析分数所定义的关联几何中的响应。位移幅度最多解释绝对响应中15%的样本外方差,且幅度匹配的配对表现出系统性差异,因此标量幅度不能组织响应。Gramian体积的闭式一阶展开产生了方向几何响应(DGR):位移在局部体积梯度上的投影,该投影联合捕获了清洁工作点、位移幅度和位移方向。绝对一阶DGR项解释了观测到的响应,样本外R²为0.838-0.969,匹配幅度排序准确率为0.864-0.963,响应符号准确率为0.909-0.989,而测试的无方向替代方法在相应评估协议下表现较弱或不稳定。预先指定的增益归一化候选V/(g_V+eps)未能通过其可预测性和清洁顺序门控。DGR使用观测到的退化状态位移,因此是一个解释性量,而非部署时的预测器:几何响应取决于表示的工作点、退化移动关联几何的程度以及移动的方向。
英文摘要
Geometric alignment scores based on Gram determinants provide a compact way to model higher-order consistency among modalities, yet how such scores respond to modality degradation is poorly understood. This paper asks whether the response of a multimodal geometric score is determined primarily by the magnitude of the perturbation-induced displacement. Using frozen cohorts from MSR-VTT (N=878) and DiDeMo (N=980), we apply controlled video blur and audio noise and analyze the response in the relational geometry on which the score is defined. Displacement magnitude explains at most 15% of the out-of-sample variance in the absolute response, and magnitude-matched pairs respond systematically differently, so scalar magnitude does not organize the response. The closed-form first-order expansion of the Gramian volume yields the Directional Geometric Response (DGR): the projection of the displacement onto the local volume gradient, which jointly captures the clean operating point, displacement magnitude, and displacement direction. The absolute first-order DGR term explains the observed response with out-of-sample R^2 of 0.838-0.969, matched-magnitude ranking accuracies of 0.864-0.963, and response-sign accuracies of 0.909-0.989, whereas the tested direction-free alternatives remain weak or unstable under the corresponding evaluation protocols. A pre-specified gain-normalization candidate, V/(g_V+eps), fails its predictability and clean-order gates. DGR uses the observed degraded-state displacement and is therefore an explanatory quantity, not a deployment-time predictor: geometric response depends on where the representation operates, how far degradation moves the relational geometry, and in which direction it moves.
CommentsSubmitted to IEEE Transactions on Multimedia (TMM). 12 pages, 6 figures, 3 tables