AI 中文总结
本文在生产级LLM智能体运行时中预注册部署了一个经破坏性验证的自动化观察器,以检测“失败看似合理”故障,通过确定性信号与LLM评判器两层流水线,机制化回归扫描并保留人类对新奇模式的关注。
AI 中文摘要
先前一项针对生产级LLM智能体运行时中静默故障的纵向研究(arXiv:2606.14589)发现,约70%的此类故障是由人类以用户身份查看产品时发现的,而数千项测试和治理检查却保持绿色通过,该研究将机制化人类视觉的部分功能视为一个开放问题。本文报告了我们对此问题的尝试。我们构建了一个自动化用户视角观察器,针对最危险的故障类别——“失败看似合理”故障,即内部错误转化为流畅、看似合理的输出呈现给用户。该观察器采用两层流水线:从事故事后分析中提炼出的五个确定性信号升级至一个LLM评判器,其裁决必须引用逐字证据,否则将被丢弃。真实基准为24个带有明确诚实边界的标注生产事故事后分析:其中24个中的16个对任何内容读取型观察器在结构上不可见,我们对此如实说明。离线状态下,确定性层实现了6/6的回归检测,且0/4的误报,每个检测器均通过破坏性验证证明其承载能力;对新模式的保留召回率为0/4。目前,它仍是一个回归引擎。部署遵循预注册流程:影子模式在首次运行中捕获并淘汰了一个系统性误报,随后26天的影子窗口运行干净,翻转标准在读取影子数据前已固定,分析协议在强制模式窗口开启前已冻结。该窗口(12个观察日)未触发任何裁决;根据预注册路径,我们将实时精确度报告为未定义,而非将平静叙述为成功。观察器自身也产生了十个静默故障,证实评判器继承了它所评判的分类体系。我们将语料库、检测器和评分卡作为可运行的基准发布:如今的机制化取代了人类的回归扫描,使人类之眼能专注于新颖性;预测仍属开放问题,但现已在冻结规则下可测量。
英文摘要
A prior longitudinal study of silent failures in a production LLM agent runtime (arXiv:2606.14589) found that about 70% were discovered by a human looking at the product as a user while thousands of tests and governance checks stayed green, and posed mechanizing part of what the human eye does as an open problem. This paper reports our attempt. We built an automated user-viewpoint observer targeting the most dangerous class, fail-plausible failure, in which an internal error becomes fluent, plausible output to the user. It is a two-layer pipeline: five deterministic signals distilled from incident postmortems escalate to an LLM judge whose verdicts must cite verbatim evidence or be discarded. Ground truth is 24 labeled production postmortems with explicit honesty boundaries: 16 of 24 are structurally invisible to any content-reading observer, and we say so. Offline, the deterministic layer achieves 6/6 regression detection with 0/4 false positives, each detector proven load-bearing by sabotage; held-out recall on novel patterns is 0/4. It is, so far, a regression engine. Deployment followed pre-registration: shadow mode caught and retired one systematic false positive on its first run, the 26-day shadow window then ran clean, flip criteria were fixed before the shadow data was read, and the analysis protocol was frozen before the enforcing-mode window opened. That window (12 observed days) fired zero verdicts; per the pre-registered path we report live precision as undefined rather than narrating quiet as success. The observer also produced ten silent failures of its own, confirming that the judge inherits the taxonomy it judges. We release the corpus, detector, and scorecard as a runnable bench: mechanization today retires the human's regression scanning so the eye can specialize in novelty; prediction remains open but is now measurable under frozen rules.
CommentsFollow-up to arXiv:2606.14589. 20 pages, 4 figures. Labeled incident corpus, detector source, sabotage-validation harness, and pre-registration texts publicly available at https://github.com/bisdom-cell/openclaw-model-bridge