发表机构
University of Maryland(马里兰大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对LLM长上下文推理易传播早期错误的问题,提出Chained RLM推理架构,通过分阶段管理上下文提升多轮迭代推理准确率,验证其相比直接LLM回答的性能增益。
AI 中文摘要
大型语言模型(LLMs)的长上下文推理通常受限于单一推理轨迹需同时探索上下文、存储中间状态、验证证据并生成最终答案,这在需要提取、计数、排序或多跳推理的任务中尤为困难,早期错误会传播至最终响应。本研究提出一种推理时架构——链式递归语言模型(Chained RLM),其中同一基础模型作为一系列新推理根被反复调用。每个推理根接收原始问题和上下文,但不继承完整对话历史,而是接收由前序推理根生成的紧凑纯文本摘要、纯文本黑板及一些持久的特定任务构件。其动机是通过将上下文拆分为部分任务而非单一大型推理响应来管理上下文,在每个分阶段计算中,中间构件可被同一模型后续的新推理检查、修正和扩展。我们描述了该系统的系统模型、交接机制、构件工作空间及评估协议,研究了在即使使用递归工具调用时,新上下文构件续接是否能比直接LLM回答带来可测量的准确率提升。
英文摘要
Long context reasoning in large language models (LLMs) is usually constrained by the fact that a single inference trajectory has to simultaneously explore the context, store intermediate state, verify evidence, and produce the final answer. This becomes particularly difficult in tasks that require extraction, counting, ordering, or multi-hop reasoning, where an early mistake can propagate until the final response. In this work, we propose Chained Recursive Language Models (Chained RLM), an inference-time architecture, in which the same underlying model is called repeatedly as a sequence of fresh reasoning roots. Each root receives the original problem and context, but does not inherit the full conversational history. Instead, it receives a compact plain-text summary, a plain-text blackboard, and some durable task-specific artifacts written by predecessor roots. The motivation is to manage the context by chopping into partial tasks rather than one large inference response; in each staged computation, intermediate artifacts can be inspected, corrected, and extended by a later fresh inference by the same model. We describe the system model, handoff mechanism, artifact workspace, and evaluation protocol for this system. We study when fresh-context artifact continuation gives a measurable gain in accuracy over direct LLM answering even with recursive tool-calling.