arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

确信错误:前沿大语言模型规则评估中的异常链崩溃

Confidently Wrong: Exception Chain Collapse in Frontier LLM Rule Evaluation

Paul Simpson, John Kozak, Lisa Doake

arXiv 2607.23386首次发表:更新:

发表机构

Aethis(Aethis)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究前沿大语言模型在规则评估中异常链崩溃问题,提出神经符号架构Aethis资格模块,通过多方面验证,将不确定性从推理边界转移到规范边界,且相关内容公开可重现。

AI 中文摘要

我们记录了前沿大语言模型中的一种失败类型——异常链崩溃,它出现在形如“A是必需的,除非B适用,除非C覆盖B”的嵌套条件规则的资格评估中。这种失败在首次观察时就会重现,但其经验表现并不稳定。我们提出了Aethis资格模块,这是一种神经符号架构,其中大语言模型根据权威来源编写规则,基于SMT的层确定性地执行这些规则。通过三个证据基础验证了该模块的有效性,包括跨四个监管领域的225个场景的受控基准测试、建筑保险的20个场景对抗扩展以及九个同行评审的LegalBench任务的外部验证。其贡献在于将不确定性从推理边界转移到规范边界,所有场景、规则编码和结果都是公开且可重现的。

英文摘要

We document a failure class in frontier large language models -- exception chain collapse -- observed in eligibility evaluation under nested conditional rules of the form "A is required UNLESS B applies, UNLESS C overrides B". The failure reproduces at first observation, but its empirical surface is unstable: between March and April 2026 several failure cells closed silently under the same model alias, with no version bump (GPT-5.4 on construction insurance moved from 96.6% to 100%, same prompt and harness). For regulated workflows, frontier-model accuracy is a moving compliance boundary that shifts without notice. We present the Aethis Eligibility Module, a neuro-symbolic architecture in which LLMs author rules from authoritative sources and an SMT-based layer executes them deterministically, consistent with the authored specification regardless of model drift, reasoning-effort defaults, or prompt format. Three evidence bases: (i) a controlled benchmark of 225 scenarios across four regulatory domains documents the pattern and, in replication, the drift that partially closed it; (ii) a 20-scenario adversarial extension on construction insurance, where the engine scores 20/20, as does one of four frontier configurations (GPT-5.4 at low reasoning effort), while the other three, including Anthropic's strongest model at evaluation time, fail the same coverage-gap edge case; (iii) external validation on nine peer-reviewed LegalBench tasks, 949 held-out cases, where the engine is significantly more accurate than all three frontier models (combined McNemar's p <= 0.003), with margins up to +41 points on the curated multi-prong tasks against the Anthropic models. The contribution is to relocate uncertainty from the inference boundary, where it is silent, to the specification boundary, where it is deliberate and audited. All scenarios, rule encodings, and results are public and reproducible.

Comments43 pages, 9 figures. Working paper v3.13.0. Benchmark data, rule encodings and per-case result artefacts: https://github.com/Aethis-ai/confidently-wrong-benchmark

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑