发表机构
Independent Researcher(独立研究者)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究智能语言模型编码工具中Claude Code的失败模式,即超时命令部分输出被误记为确认结果并传播误报,揭示其观察与持久性 conflation 的机制,指出对依赖会话连续性的工作流程有影响。
AI 中文摘要
智能语言模型编码工具将长时间会话历史压缩为压缩摘要,后续会话将其作为基本事实继承。本文记录了Claude Code中的一种失败模式,即超时命令(退出代码143)的部分标准输出在压缩摘要中被记录为确认结果,在未重新验证的情况下跨会话和模型版本传播误报。其潜在机制是观察与持久性的 conflation,终端中出现的信息被视为等同于写入持久存储的信息。这一发现扩展了先前关于语言模型作为评判分级中的非确定性的工作中报告的语言模型自我评估失败的分析,表明智能工具在报告自身操作结果时也存在类似的可靠性缺陷。该失败对任何依赖智能会话连续性进行数据处理、科学计算或多步自动化的工作流程都有直接影响。
英文摘要
Agentic LLM coding tools compress long session histories into compaction summaries that subsequent sessions inherit as ground truth. This paper documents a failure mode in Claude Code where partial standard output from timed-out commands (exit code 143) is recorded in compaction summaries as confirmed results, propagating false positives across sessions and model versions without re-verification. The underlying mechanism is a conflation of observation and persistence, where information that appeared in the terminal is treated as equivalent to information written to durable storage. This finding extends the analysis of LLM self-evaluation failures reported in prior work on non-determinism in LLM-as-judge grading by showing that agentic tools exhibit analogous reliability deficits when reporting on their own operational outcomes. The failure has direct implications for any workflow that relies on agentic session continuity for data processing, scientific computation, or multi-step automation.
Comments8 pages, companion to arXiv:2606.26185