arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

多模态 Transformer 中的细化对称性

Refinement Symmetry in Multimodal Transformers

Yuhao Du, Shunian Chen

arXiv 2609.32669首次发表:更新:

AI 中文总结

本文提出多模态 Transformer 中的细化对称性原理,通过度量加权保持拆分表示贡献,在 Qwen2.5-Omni-7B 上提升跨分区鲁棒性,显著减少注意力误差并提高准确率。

AI 中文摘要

注意力权重依赖于 token 数量,而 token 数量会随信号的表示方式而变化。我们研究细化对称性:在保持内容、位置、可见上下文和总质量不变的情况下,拆分表示应保持其贡献。基于比例注意力和正交注意力,我们证明拆分不变性迫使局部质量因子对于任何固定的正注意力核必须是线性的,前提是该因子是非递减的。对于改变后的表示,物理耦合通过将特征变化与权重重新分配分离来限制注意力误差。在 Qwen2.5-Omni-7B 中,将一半视觉 token 重复三倍,在标准注意力下改变了 3,586 个 MVBench 答案中的 255 个;而度量加权在匹配可见性下保留了所有答案。在自然帧重采样下,它减少了分布漂移。在冻结视频编码的两倍合并中,五次种子评估显示,与全局计数加权(平均组质量)相比,全分区正确准确率提高了 1.04 个百分点,与标准注意力相比提高了 0.93 个百分点。相对于全局计数的优势在 WorldSense 上也成立,但取决于压缩预算。结果是一个表示原则,在跨分区的鲁棒性方面具有可衡量的收益。

英文摘要

Attention weights depend on token counts, which change with the representation of a signal. We study refinement symmetry: splitting a representation while preserving content, position, visible context, and total mass should preserve its contribution. Building on proportional and quadrature attention, we show that split invariance forces the local mass factor to be linear for any fixed positive attention kernel, provided that factor is nondecreasing. For changed representations, a physical coupling bounds attention error by separating feature change from weight reallocation. In Qwen2.5-Omni-7B, duplicating half the visual tokens threefold changes 255 of 3,586 MVBench answers under standard attention; measure weighting preserves every answer under matched visibility. Under natural frame resampling, it reduces distributional drift. At twofold merging of a frozen video encoding, a five-seed evaluation shows an all-partition-correct accuracy gain of 1.04 percentage points over global count weighting (average group mass) and 0.93 points over standard attention. The advantage over global count also holds on WorldSense but depends on the compression budget. The result is a representation principle with a measured benefit in robustness across partitions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑