佩内洛普:用于高效结构化推理的局部潜在循环
Penelope: Localized Latent Recurrence for Efficient Structured Reasoning
浏览论文内容
中文总结 AI 辅助
研究复杂结构化推理任务计算问题,提出佩内洛普框架,通过局部潜在循环实现高效推理,将可见推理转移到内部循环路径,在开源基准实验中降低推理延迟,实现准确性-效率权衡。
中文摘要 AI 辅助
复杂的结构化推理任务通常需要额外的计算,而当前语言模型主要通过增加参数规模或序列化中间步骤(如思维链(CoT)令牌)来实现。前者增加了训练和部署成本,后者将推理计算与自回归输出长度联系起来。我们引入了佩内洛普,这是一个用于预训练的仅解码器变压器的高效潜在推理框架,它将循环计算定位到选定的解码器区间。较低的解码器前缀被评估一次以构建问题条件边界内存,然后在答案生成之前通过时间调制的GRU动态和循环读出状态进行迭代细化。渐进式CoT到潜在课程将可见推理转移到这个内部循环路径中,允许在潜在空间中分配额外的计算,而无需重复执行完整的解码器或生成长的中间痕迹。在开源结构化推理基准上的实验表明,在验证选择的潜在预算下,佩内洛普相对于已建立的潜在推理模型具有有竞争力的准确性,同时降低了测量的推理延迟。这些结果表明,潜在细化可以定位到一个狭窄的解码器区间,减少重复的全解码器执行,而不产生长的可见推理痕迹,并为仅解码器变压器模型提供了实际的准确性-效率权衡。
英文摘要
Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens. The former raises training and deployment costs, while the latter ties reasoning computation to autoregressive output length. We introduce Penelope, an efficient latent-reasoning framework for pretrained decoder-only Transformers that localizes recurrent computation to a selected decoder interval. The lower decoder prefix is evaluated once to construct a problem-conditioned boundary memory, which is then iteratively refined through time-modulated GRU dynamics and recurrent readout states before answer generation. A progressive CoT-to-latent curriculum transfers visible reasoning into this internal recurrent path, allowing additional computation to be allocated in latent space without repeatedly executing the complete decoder or generating a long intermediate trace. Experiments on open-source structured-reasoning benchmarks show that, at validation-selected latent budgets, Penelope attains competitive accuracy relative to established latent-reasoning models while reducing measured inference latency. These results show that latent refinement can be localized to a narrow decoder interval, reducing repeated full-decoder execution without generating a long visible reasoning trace and providing a practical accuracy-efficiency tradeoff for decoder-only Transformer models.