arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

RIG-RoPE:具有时长感知时间坐标的关系与实例门控旋转位置编码

RIG-RoPE: Relation-Stratified Multimodal Attention with Instance-Local Rotary Geometry and Representation-Aware Traversal Coordinates

Donggen Li

arXiv 2608.05154首次发表:更新:

发表机构

Sichuan University(四川大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对多模态场景中静态多维位置分配的局限,提出RIG-RoPE机制,通过关系与实例门控及时长感知时间坐标优化位置编码,未增加可学习参数,确立了相关公式与验证路径。

AI 中文摘要

旋转位置编码(RoPE)是现代语言模型的核心组件,已通过多模态变体(如多模态RoPE(M-RoPE))扩展至多模态大语言模型,该变体将位置通道拆分为时间、高度和宽度子空间。本报告指出了交错多模态场景中静态多维位置分配的两个局限:其一,高度/宽度旋转可能应用于空间位移非明确定义几何对象的标记对,产生跨模态及跨实例空间干扰;其二,时间坐标常被视为等步长计数器,导致文本标记、图像块、视频段虽信息密度不同,却能推进相当量的时间相位。我们提出RIG-RoPE,一种具有时长感知时间坐标的关系与实例门控RoPE机制。RIG-RoPE为每个标记增加模态指示器、视觉实例标识符及标量信息时长坐标,仅对来自同一视觉实例的查询-键对启用高度/宽度旋转,否则将未知空间位移边缘化而非设为零;时间旋转采用插值累积块时长:文本标记消耗单位时长,图像使用维度感知对数空间尺度,视频则在有效帧上进一步应用对数时间扩展。我们提供了避免普通跨实例空间旋转的规范不变性论证、共享高度/宽度子空间下静态标识符的不可能性结果,以及反对等步长多模态时间的时长一致性论证。RIG-RoPE未增加可学习参数,可在分块注意力核中实现,每个标记仅需恒定额外元数据。本初步报告确立了公式与验证路径,未声称经验优越性。

英文摘要

Multimodal rotary positional encodings apply temporal, height, and width phases to interleaved text, image, and video tokens. This creates two ambiguities: cross-instance spatial displacement depends on preprocessing chart choices unless registration is declared, and scalar advance across visual blocks is often inherited from coordinate extrema rather than defined at the representation level. We introduce RIG-RoPE, combining instance-local rotary geometry, relation-stratified attention, and representation-aware traversal coordinates. RIG-RoPE normalizes relation-homogeneous scores separately, allocates mass with a common H/W-neutral LogSumExp statistic, and uses traversal extent that is additive over ordered slices and sublinear over parallel spatial scale. Text advances by unit increments, image patches are simultaneous, and video accumulates over tokenizer temporal tokens. In a matched, inference-only Qwen2-VL-2B checkpoint experiment, native and RIG text-only paths were exactly equal. RIG was exactly invariant to a whole-chart single-image translation and to translating only the second instance of an unregistered image pair. Native attention remained sensitive to the latter, while an H/W-collapse control confirmed that RIG retained same-instance spatial effects; visual embeddings and all parameters were unchanged. Across three seeds of a frozen tiny task, RIG also had zero clean-to-Gauge logit change, whereas the raw-H/W baseline changed in every seed. Gauge-accuracy differences were +2/72, 0, and 0, failing the preregistered stability gate. These results support the specified activation and Gauge mechanisms, not stable task improvement, universality, or empirical superiority.

Comments26 pages, 2 figures, 4 tables. This revision adds matched, inference-only Qwen2-VL-2B native/RIG L1 mechanism evidence: exact text-only equivalence, exact RIG Gauge invariance, native Gauge sensitivity, and retained same-instance spatial effects. Task-level Qwen benchmarks, training gains, stable task improvement, universality, and empirical superiority remain unestablished

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑