AI 中文总结
本研究探究LLM在低可控失效概率工作流中,解释性参与度随失效渐近稀有性的变化,发现触发结构是关键调节变量,不同模型的异常识别与参与表现存在差异。
AI 中文摘要
现有关于大语言模型(LLM)在异常条件下行为的研究,主要关注模型是否能察觉异常。本文提出更聚焦的问题:当模型处于失效概率低且可控的工作流中时,其解释性参与度(长度、特异性、自我报告置信度)是否会随失效的渐近稀有性提升而变化?我们在三个开源权重模型(qwen3:8b、llama3.1:8b、mistral:7b)上构建了本地零成本测试工具,开展重复工具调用任务:其中一次调用的失效概率p在0.2至0.0001的8个区间内变化,同时设置从即时提示到无提示的5种触发条件。我们假设:失效越稀有,模型参与度会先上升,随后在可检测阈值附近崩溃。跨条件汇总后,该假设不成立:解释长度呈平稳单调下降趋势;但按条件拆分后结果反转:在要求模型立即解释每一次失效的immediate_forced条件下,预测的上升趋势得到验证,随后进入平台期而非崩溃——p=0.05时解释长度峰值达28.4词,在最稀有率下稳定在17.4-19.0词,置信度从约53%不均等地升至70%-90%区间;在将解释批量到运行结束的grouped_runs条件下,未出现崩溃;在无提示的passive_unprompted条件下,总幅度为基线假象,但修复日志缺口后发现真实的模型特异性自我监控:llama3.1:8b会主动在无提示下提供结构化置信度报告,随试验累积有时会降低自身置信度,另外两个模型仅作为模板出现一次。触发结构是崩溃可观测性的重要调节变量。配套的保证失效运行(72个单元,回填随机抽样未出现真实失效的区间)显示,模型识别异常的能力与识别后的参与度存在差异。局限性:离散率点无法捕捉区间内的行为,是未来研究方向。
英文摘要
Prior work on LLM behavior under anomalous conditions asks whether a model notices anomalies. We ask a narrower question: once a model sits in a workflow with a low, controllable failure rate, does its explanatory engagement - length, specificity, self-reported confidence - change as failure grows asymptotically rarer? We built a local, zero-cost harness on three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b) running a repeated tool-call task where one call fails at probability p, swept across eight rates from 0.2 to 0.0001, under five elicitation conditions from immediate prompting to none. We hypothesized a rise in engagement as failures grew rarer, then a collapse near a detectability threshold. Pooled across conditions this appeared false: length fell in a flat, monotonic pattern. Splitting by condition overturned that. Under immediate_forced, where the model must explain every failure instantly, the predicted rise is confirmed but followed by a plateau, not a collapse: length peaks at 28.4 words at p=0.05, settles to 17.4-19.0 words at the rarest rates, and confidence rises unevenly from about 53% to the 70s-90s. Under grouped_runs, explanation batched to run-end, no collapse appears. Under passive_unprompted, aggregate magnitude is a floor artifact, but a recovered logging gap revealed real, model-specific self-monitoring: llama3.1:8b volunteers structured confidence reports unprompted, sometimes eroding its own confidence as trials accumulate; the other two do so only once, as boilerplate. Elicitation structure is a first-class moderator of collapse observability. A companion guaranteed-failure run (72 cells, backfilling rates where random sampling gave zero real failures) shows models differ in whether they recognize an anomaly, distinct from engagement once recognized. Limitation: discrete rate points cannot capture behavior between them, a direction for future work.
Comments11 figures. Elicitation-condition sweep across three open-weight models (qwen3:8b, llama3.1:8b, mistral:7b); pipeline scripts and experimental data available upon reasonable request