AI 中文总结
ROTE基准通过LZW复杂度控制的符号序列,评估神经架构的精确记忆能力,揭示记忆质量、滚动稳定性与计算开销之间的权衡。
AI 中文摘要
我们引入了ROTE(精确记忆的滚动测试),一个用于评估神经架构符号记忆的基准测试协议。我们通过使用复杂度由Lempel-Ziv-Welch(LZW)压缩调控的序列,研究神经序列模型中的记忆和符号规则的扩展。在ROTE下,每个架构被训练为相同的有限上下文条件预测器,并使用教师强制的单步预测以及闭环滚动对保留符号进行评估。遵循共享的预测和滚动评估例程,该基准评估了门控循环、最小循环、基于注意力和混合循环注意力模型,并保留其原生计算特性。除了标准预测指标外,基准还报告了归一化字符串距离、训练时间、内存使用和参数数量,跨越LZW复杂度扫描。该研究建立了算法序列复杂度与神经架构记忆能力之间的联系,揭示了记忆质量、滚动稳定性和计算开销之间的权衡。我们的软件和可复现实验代码可从该https URL获取。
英文摘要
We introduce ROTE (RollOut Testing of Exact memorization), a benchmarking protocol for evaluating symbolic memorization of neural architectures. We study memorization and the extension of symbolic rules in neural sequence models by using sequences whose complexity is regulated by Lempel--Ziv--Welch (LZW) compression. Under ROTE, each architecture is trained as the same finite-context conditional predictor and is evaluated using teacher-forced one-step prediction as well as closed-loop rollout on the withheld symbols. Following a shared prediction-and-rollout evaluation routine, the benchmark evaluates gated recurrent, minimal recurrent, attention-based, and hybrid recurrent-attention models with their native computational characteristics preserved. Beyond standard predictive metrics, the benchmark reports normalized string distances, training time, memory usage, and parameter count across an LZW-complexity sweep. The study establishes a connection between the complexity of algorithmic sequences and the memorization capacity of neural architectures, revealing the trade-offs involving memorization quality, rollout stability, and computational expense. Our software and reproducible experimental code can be obtained from https://github.com/nla-group/rote.