arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

TraceRelay:基于滚动轨迹的注意力对齐循环架构

TraceRelay: Attention-Aligned Recurrence over Rolling Traces

Sungwoo Goo, Hwi-yeol Yun, Sangkeun Jung

arXiv 2610.11743首次发表:更新:

发表机构

College of Pharmacy, Chungnam National University; Department of Computer Science & Engineering, Chungnam National University(忠南大学药学院; 忠南大学计算机科学与工程系)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出TraceRelay架构,在Equal Repeats等任务上验证其循环阶段继承等设计的有效性,发现循环轨迹维度对不同任务的性能影响存在差异,需进一步研究循环表示大小的选择规则。

AI 中文摘要

我们提出了TraceRelay,这是一种注意力对齐的循环架构,可在低维轨迹的滚动序列上分布持久表示。局部右向注意力从低层表示中形成增量;传递会延迟,直到所有被关注的输入都处于因果过去。固定的加法阶段循环会累积延迟的增量,而左向注意力则读取由此产生的残差增强流。步长前缀和支持并行预填充和有界缓冲区续接。我们在Equal Repeats、有界Dyck闭合类型预测和因果最频繁生成任务上,对36次小模型运行进行了研究,每个设置使用3个随机种子。在训练长度为256时,具有循环阶段继承的Equal Repeats模型达到98.81%-99.69%的准确率,而未继承的单独训练变体仅为50.73%-51.63%,尽管后者接受了更多更新。在超过训练范围的长度下,准确率会急剧下降。在评估的最长长度下,中间层循环轨迹维度更多的模型在Dyck任务上表现更好(长度4096时,接近准确率为76.34%,而维度更少的模型为55.49%),而轨迹维度更少的模型在五符号最频繁生成任务上表现更好(长度1024时,精确生成为70.74%,而维度更多的模型为55.60%)。这些对比案例表明,需要进一步研究如何针对不同任务选择循环表示的大小,目前尚未在任务或模型配置间建立通用规则。

英文摘要

We present TraceRelay, an attention-aligned recurrent architecture that distributes persistent representations over a rolling sequence of low-dimensional traces. Local right looking attention forms increments from lower-layer representations; delivery is delayed until all attended inputs are in the causal past. A fixed additive phase recurrence accumulates the delayed increments, and left-looking attention reads the resulting residual augmented stream. A stride-wise prefix sum supports parallel prefill and bounded-buffer continuation. We study 36 small-model runs on Equal Repeats, bounded Dyck closing-type prediction, and causal Most-Freq generation, using three seeds per setting. At trained length 256, Equal Repeats models with recurrent phase inheritance reach 98.81-99.69% accuracy versus 50.73-51.63% for separately trained variants without inheritance, despite the latter receiving more updates. Accuracy drops sharply at lengths beyond the training range. At the longest evaluated lengths, models with more dimensions in the middle layer's recurrent traces perform better on Dyck (76.34% versus 55.49% close accuracy at length 4096), whereas models with fewer trace dimensions perform better on five-symbol Most-Freq (70.74% versus 55.60% exact generation at length 1024). These contrasting cases motivate further study of how the size of recurrent representations should be chosen for different tasks, without establishing a general rule across tasks or model configurations.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑