AI 中文总结
该研究提出OAttention与O-Closure,通过活跃存在系数实现零向量标记的惰性属性,经测试在TabPFN v3零微调中可小幅提升回归性能,验证了其精确性与路径兼容性。
AI 中文摘要
注意力掩码是关系级控制:它们指定哪些查询-源对可以交互。它们不提供在注意力边界处不参与的、携带表示的标记状态。我们为每个标记的隐藏载体\boldsymbol{h}_i分配一个活跃存在系数\boldsymbol{p}_i = \boldsymbol{\rVert h_i \rVert}^2 / (\tau + \boldsymbol{\rVert h_i \rVert}^2)。该系数有两个作用:它控制标记i发出的信息,并决定标记i进入与其他标记共享的计算的质量。OAttention是该规则的支持耦合注意力实现,它通过\boldsymbol{p}_i控制接收方输出,并在注意力分子和划分中均通过\boldsymbol{p}_j对源j进行加权,同时保留标准分数、可见性关系、指数竞争和值聚合。这使得零向量标记成为零元素,并产生精确的空接收方、空源插入、自注意力插入和空支持属性。相同的标记级存在产生局部O组件(OFFN、ONorm和OInject)、存在加权的OStandardize、O-Closure定律\boldsymbol{M(H \bigoplus 0) = M(H) \bigoplus 0},以及通过残差和组合闭合得到的OTransformer。规范算子通过契约测试和GPU评估进行验证。在克隆的预训练TabPFN v3回归器的零微调改造中,经过校准的隐藏载体OAttention和Full-O变体在18个匹配的数据集-种子案例上,平均RMSE分别变化+0.088%和+0.177%。双模块消融实验表明,单独的OAttention无法通过普通宿主组件保留NULL状态,而OTransformer路径可以。这些是对精确性、活跃路径兼容性和组合必要性的范围测试;它们并未确立通用的无损失、任意宿主闭合、对原点的学习吸引力或缺失值的通用语义。
英文摘要
Attention masks are relation-level controls: they specify which query--source pairs may interact. They do not provide a representation-carried token state that is non-participating at the attention boundary. We assign each token hidden carrier \(h_i\) an active-presence coefficient \(p_i=\lVert h_i\rVert^2/(τ+\lVert h_i\rVert^2)\). The same coefficient has two roles: it gates information emitted by token \(i\), and it determines the mass with which token \(i\) enters computations shared with other tokens. OAttention is the support-coupled attention realization of this rule. It gates the receiver output by \(p_i\) and weights source \(j\) by \(p_j\) in both the attention numerator and partition, while retaining the standard score, visibility relation, exponential competition, and value aggregation. This makes the zero-vector token a zero element and yields exact null-receiver, null-source insertion, self-attention insertion, and empty-support properties. The same token-level presence gives local O-components (OFFN, ONorm, and OInject), presence-weighted OStandardize, the O-Closure law \(M(H\oplus0)=M(H)\oplus0\), and an OTransformer by residual and compositional closure. The canonical operator is checked by contract tests and a GPU evaluation. In a zero-fine-tuning retrofit of a cloned pretrained TabPFN v3 regressor, calibrated hidden-carrier OAttention and Full-O variants change mean RMSE by $+0.088\%$ and $+0.177\%$, respectively, over 18 matched dataset--seed cases. A two-block ablation shows that OAttention alone does not preserve a NULL state through ordinary host components, whereas the OTransformer path does. These are scoped tests of exactness, active-path compatibility, and compositional necessity; they do not establish universal no-loss, arbitrary-host closure, learned attraction to the origin, or a general semantics for missing values.