AI 中文总结
本文报告了数月LLM辅助开发中持久代理记忆的运维经验,提出SIx Harness框架,通过可靠性工程实现记忆子系统的高可用,并提炼七项设计原则。
AI 中文摘要
LLM编码代理正从任务规模跨越到项目规模:单一辅助工作持续数月,经历反复的上下文压缩,处理远超任何上下文窗口的代码库。我们报告了此类工作(自2026年1月起在单一连续Claude Code会话线下运行的研究项目,驱动一个633,000行代码库,其记忆子系统自2026年7月起持续插桩)的运维经验,并描述了SIx Harness,即从保持其一致性中产生的开源记忆与连续性基础设施。该框架结合了每项目长期记忆(基于SQLite的混合词法-向量检索,完全本地化)、精度门控的上下文注入、针对决策和死胡同的反复发存储、为在压缩中存活而设计的约定,以及我们认为在运维实践中缺失的部分:记忆子系统本身的可靠性工程——具有区分故障模式的会话启动健康门、设计为确保任何列举的故障模式都无法不被记录的心跳遥测,以及借鉴自SRE实践的告警疲劳预算。我们展示了我们认为的首个在生产开发使用中持续数月的持久代理记忆子系统的插桩运维记录:78,933次钩子调用;85次记录故障,无一静默:其中84次发生在子系统前两周,此后1次,最后20天无故障;一个注入层,其十天精度仪表在完整分母下显示零误报;以及三个从仪表读数追溯到结构性修复的生产事件。从该记录中,我们提炼出七项设计原则,坦率陈述我们的局限性(N=1,无对照组,自我报告),并发布一个带标签的预注册消融协议,任何团队都可以使用已发布的MIT许可工具包运行。
英文摘要
LLM coding agents are crossing from task-scale to project-scale: single assisted efforts that run for months, across repeated context compactions, on codebases far larger than any context window. We report operational experience from one such effort (a research project under a single continuous Claude Code session line since January 2026, driving a 633,000-line codebase, with its memory subsystem continuously instrumented since July 2026) and describe SIx Harness, the open-source memory and continuity infrastructure that emerged from keeping it coherent. The harness combines per-project long-term memory (hybrid lexical-vector retrieval over SQLite, fully local), precision-gated context injection, anti-recurrence stores for decisions and dead-ends, conventions engineered to survive compaction, and, the part we argue is missing from operational practice, reliability engineering for the memory subsystem itself: a session-start health gate with discriminated failure modes, heartbeat telemetry designed so that no enumerated failure mode can pass unrecorded, and alert-fatigue budgeting borrowed from SRE practice. We present what we believe is the first months-scale instrumented operational record of a persistent agent-memory subsystem in production development use: 78,933 hook invocations; 85 recorded failures, none silent: 84 in the subsystem's first three weeks, one since, none in the final 20 days; an injection layer whose ten-day precision instrument shows zero false fires against an intact denominator; and three production incidents traced from instrument reading to structural fix. From the record we distill seven design principles, state our limitations plainly (N=1, no control arm, self-reported), and publish a tagged pre-registered ablation protocol that any team can run with the released MIT-licensed kit.