arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

可信AI的真实性:表达怀疑、来源与信念修正作为可工程化的立场

Truth for Believable AI: Expressed Doubt, Provenance, and Belief Revision as an Engineerable Stance

Sebastian Cochinescu

arXiv 2609.26035首次发表:更新:

发表机构

University of Bucharest(布加勒斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究在固定语言模型上实现表达怀疑、来源与信念修正的行为层,通过合成和真实模型评估,发现令牌概率门控失败而一致性门控有效,需扩展事实库后评估。

AI 中文摘要

对话代理通常以统一自信的语气表达答案。我们测试了表达不确定性、来源感知的断言和显式信念修正是否可以作为一个行为层在固定语言模型上实现;我们不测试可信度或信任。该层结合了三种认知状态:每条主张的置信度和类型化来源、来源门控的表达规则,以及具有可审计确认和对虚假修正部分抵抗的持久修正存储。我们使用合成模型和Qwen2.5-0.5B-Instruct在一个构建的、机械评分的多会话基准上对其进行评估。合成工具通过了所有五项检查。在真实模型上,确认健全性(一种构造保证)在100%的情况下成立,并且真实修正比虚假修正更常被接受(对已有信念为0.44对0.15;包括规则接受的未持有事实修正为0.875对0.420),但预先指定的表达保真度、矛盾分离和来源边际失败。一项公开的事后分析表明,基于平均答案令牌概率门控的表达在端到端上正确性排名低于随机水平(AUC 0.41,会话聚类),而基于采样一致性门控则具有区分能力(AUC 0.66)。从这一发现中选择的一致性门控配置在单独承诺的协议下评估,满足了会话级操作和能力等价标准,并在重新绘制的会话集上复现。操作结果依赖于选择,当不确定性在60个事实上聚类时,两个标准仍未解决。支持的结论仅限于构造审计保证、依赖存储的部分修正区分,以及令牌概率门控的基准和模型特定失败;在人类评估之前需要扩展事实库。

英文摘要

Conversational agents often express answers in a uniformly confident register. We test whether expressed uncertainty, provenance-aware assertion, and explicit belief revision can be implemented as a behavior layer over a fixed language model; we do not test believability or trust. The layer combines three epistemic states, per-claim confidence and typed provenance, a provenance-gated expression rule, and a persistent revision store with auditable acknowledgments and partial resistance to false corrections. We evaluate it on a constructed, mechanically scored multi-session benchmark using a synthetic model and Qwen2.5-0.5B-Instruct. The synthetic instrument passes all five checks. On the real model, acknowledgment soundness, a by-construction guarantee, holds in 100% of cases, and true corrections are accepted more often than false ones (0.44 vs. 0.15 on held beliefs; 0.875 vs. 0.420 including rule-accepted corrections of unheld facts), but the pre-specified expression-fidelity, contradiction-separation, and provenance margins fail. A disclosed post hoc analysis shows that expression gated on mean answer-token probability ranks correctness below chance end to end (AUC 0.41, conversation-clustered), whereas gating on sampling consistency discriminates (AUC 0.66). A consistency-gated configuration selected from this finding and evaluated under a separately committed protocol meets the conversation-level manipulation and capability-equivalence criteria and replicates on a redrawn conversation set. The manipulation result is selection-dependent, and both criteria remain unresolved when uncertainty is clustered over the 60 facts. The supported conclusions are limited to the by-construction audit guarantee, store-dependent partial correction discrimination, and a benchmark- and model-specific failure of token-probability gating; scaling the fact base is required before human evaluation.

Comments17 pages, 4 figures, 3 tables. Companion framework paper: arXiv:2607.15883. Code, benchmark, cached model outputs, and result files archived at doi:10.5281/zenodo.21462986 (code and results) and doi:10.5281/zenodo.21462988 (benchmark dataset)

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑