AI 中文总结
该研究测试LLM能否通过196次依赖工具调用携带精确中间状态,发现gpt-oss-120b多数运行成功,关键因素为保留自身推理及投票消除工作模型算术错误。
AI 中文摘要
长程任务在大语言模型(LLM)评估中仍不常见,原因在于:当每一步都依赖前一步时,单独来看表现优异的每步准确率会灾难性衰减,因为错误会级联,且端到端失败概率会随长度急剧增长。现有智能体基准报告了端到端成功,但将这种状态跟踪难度与指令解释混淆,未设置隔离该难度的对照组,且易受幻觉最终答案等捷径影响,因此无法说明长运行失败的原因。LLM是否能在多次工具调用中携带精确中间状态本身也未得到充分验证。我们通过让模型逐步计算加密哈希MD5来对此进行清晰测试:这是一个在64轮中包含196次依赖工具调用的序列,期间它在自身上下文内从一次调用到下一次调用携带四个32位字(a,b,c,d)。解释工作十分简单,且由于我们从头实现了MD5(RFC 1321),我们将每一次调用与真实轨迹对齐并逐位检查摘要,因此任何失败都纯粹是簿记问题。gpt-oss-120b是一个每token仅约55亿个激活参数的混合专家模型,在温度为0且使用短固定提示的情况下,它在所有196次调用中携带了完整状态,并在多数完成的运行中返回了正确摘要。在最强设置中,我们将每个基础工具替换为第二个LLM,因此驱动模型和工作模型从头计算整个哈希,循环中没有精确算术预言机。成功取决于两个因素且二者均不改变权重:每一轮在其上下文中保留模型自身的推理,以及对支持思考的工作模型进行投票以消除其模运算错误。我们按来源定位残留失败,将状态携带与算术及服务问题分离。
英文摘要
Long-horizon tasks remain uncommon in large language model (LLM) evaluation, and for a reason: when each step depends on the last, per-step accuracy that looks excellent in isolation decays catastrophically, as errors cascade and the end-to-end failure probability grows sharply with length. Existing agentic benchmarks report end-to-end success but confound this state-tracking difficulty with instruction interpretation, give no control group that isolates it, and are vulnerable to shortcuts such as a hallucinated final answer, so they cannot say why a long run fails. Whether an LLM can carry exact intermediate state across many tool calls at all is itself not well established. We test this cleanly by having the model compute a cryptographic hash, MD5, step by step: a sequence of $196$ dependent tool calls over $64$ rounds while it carries four $32$-bit words $(a,b,c,d)$ in its own context from one call to the next. Interpretation is trivial and, because we implement MD5 from scratch (RFC~1321), we align every call to the ground-truth trace and check the digest to the bit, so any failure is pure bookkeeping. gpt-oss-120b, a mixture-of-experts model with only $\sim$5.5B active parameters per token, at temperature $0$ with a short fixed prompt, carries the full state across all $196$ calls and returns the correct digest on a majority of completed runs. In the strongest setting we replace every primitive tool with a second LLM, so a driver and a worker compute the whole hash from scratch with no exact-arithmetic oracle in the loop. Two ingredients decide success and neither changes the weights: keeping the model's own reasoning in its context each turn, and voting over a thinking-enabled worker to remove its modular-arithmetic slips. We localize the residual failures by origin, separating state-carrying from arithmetic and from serving.
Comments6 Pages, 2 figures