随机递归模型
Random Recursive Models
浏览论文内容
中文总结 AI 辅助
随机递归模型通过随机采样层池中的层进行递归计算,实现灵活层复用,在推理任务中以更少参数匹配或超越基线,并支持深度变化和测试时扩展。
中文摘要 AI 辅助
递归模型通过参数复用创造计算深度,为增加模型规模提供了一种参数高效的替代方案。然而,大多数递归模型反复应用一个学习到的变换或一组预设的变换序列,将计算限制在固定的层顺序上。我们引入了随机递归模型(RRM),该模型维护一个包含 $L$ 个学习层的池,并通过为每个样本和步骤独立地有放回地采样一个层来执行 $T$ 个递归步骤。这实现了灵活的层复用,同时保留了递归的参数效率。我们在具有挑战性的推理任务上评估了 RRM,其性能匹配或超过基线,且通常参数减少 50-75%。RRM 可以在推理时改变其深度,包括超过训练时所见深度,而无需重新训练或添加参数,从而改善了受益于更深迭代计算的任务。RRM 还支持蒙特卡洛推理和概率性测试时扩展,两者均无需重新训练即可提升性能。这些见解可能为神经网络架构设计开辟新的方向。
英文摘要
Recursive models create computational depth through parameter reuse, offering a parameter-efficient alternative to increasing model size. However, most recursive models repeatedly apply one learned transformation or a prescribed sequence of transformations, restricting computation to a fixed layer order. We introduce the Random Recursive Model (RRM), which maintains a pool of $L$ learned layers and performs $T$ recursive steps by sampling one layer independently with replacement for each example and step. This enables flexible layer reuse while retaining the parameter efficiency of recurrence. We evaluate RRM on challenging reasoning tasks, where it matches or exceeds the baselines, often with 50-75 % fewer parameters. RRM can vary its depth at inference, including beyond that seen during training, without retraining or adding parameters, improving tasks that benefit from deeper iterative computation. RRM also supports Monte Carlo inference and probabilistic test-time scaling, both of which improve performance without retraining. These insights may open new directions in neural network architecture design.
发表机构
- Mila – Quebec AI Institute(Mila – 魁北克人工智能研究所)
- Université de Montréal(蒙特利尔大学)
- Concordia University(康考迪亚大学)
机构由 AI 辅助整理,请以论文原文为准。