面向延迟约束多模态令牌通信的几何跨模态令牌选择
Geometric Cross-Modal Token Selection for Latency-Constrained Multimodal Token Communication
浏览论文内容
中文总结 AI 辅助
该研究针对延迟约束多模态令牌通信问题,提出基于几何的跨模态令牌选择框架及IBS、R-IBS策略,在VQA和AVQA任务上实现了显著的准确率提升。
中文摘要 AI 辅助
本文提出一种基于几何的联合跨模态令牌选择框架,用于延迟约束多模态令牌通信。为捕捉跨模态令牌依赖关系,我们利用交叉注意力机制将模态特定令牌投影到共享的查询-键空间,其中令牌数量最少的模态作为锚定模态,其余模态作为非锚定模态。受胚芽模型(germ-grain models)启发,我们定义了角度距离度量,并围绕锚定查询构建语义颗粒区域。基于这种几何表示,我们识别出该空间中多个锚定查询共享的跨模态证据,并开发了基于交集的令牌选择(IBS)策略,该策略优先选择键被多个颗粒区域覆盖的非锚定令牌。我们进一步开发了一种名为鲁棒IBS(R-IBS)的感知擦除扩展方法,用于令牌级擦除信道,采用预期角度距离公式。在IBS和R-IBS中,颗粒区域在延迟约束下针对单个查询进行优化,使用块坐标下降和低复杂度贪心算法。仿真验证了IBS和R-IBS在延迟约束及令牌级擦除信道下的有效性,与现有令牌选择基线方法相比,它们在视觉问答(VQA)和视听问答(AVQA)任务上分别实现了高达31.6%和29.2%的任务准确率提升。
英文摘要
This paper proposes a geometry-based joint cross-modal token selection framework for latency-constrained multimodal token communications. To capture cross-modal token dependencies, we leverage the cross-attention mechanism to project modality-specific tokens into a shared query-key space, where the modality with the fewest tokens serves as the anchor modality and the others as non-anchor modalities. Inspired by germ-grain models, we define an angular-distance metric and construct semantic grain regions around anchor queries. Based on this geometric representation, we identify cross-modal evidence shared across multiple anchor queries in this space and develop an intersection-based token selection (IBS) strategy that prioritizes non-anchor tokens whose key are covered by multiple grain regions. We further develop an erasure-aware extension, termed robust-IBS (R-IBS), for token-wise erasure channels using an expected angular-distance formulation. In both IBS and R-IBS, the grain regions are optimized for individual queries under a latency constraint, using block coordinate descent and a low-complexity greedy algorithm. Simulations corroborate the effectiveness of IBS and R-IBS under latency-constrained and token-wise erasure channels, achieving up to 31.6% and 29.2% task accuracy gains on visual question answering (VQA) and audio-visual question answering (AVQA) tasks, respectively, over existing token selection baselines.