arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

通过构建实现零幻觉:用于可信企业人工智能的幻觉感知分层监督

Zero Hallucination, by Construction: Hallucination-Aware Layered Oversight for Trustworthy Enterprise AI

Bogdan Raduta, Horia Velicu, Alexandru Preda, Serban Chiricescu

arXiv 2607.17883首次发表:更新:

AI 中文总结

研究企业AI因幻觉难以被信任的问题时,提出HALO架构,通过六层防御将幻觉视为可控制故障模式,详细介绍各层并关注基于证据的置信度,以实现可信企业AI,在索赔提取工作负载上进行了架构说明。

AI 中文摘要

企业不会部署他们不信任的人工智能代理,而最常被提及的不信任原因是幻觉,即自信、流畅但不正确的输出。常见的应对方法是等待一个不会产生幻觉的模型,而本文认为这是错误的目标。大语言模型天生就能够生成无根据的文本,增加规模也无法消除这种可能性。本文将目标重新定义为:“零幻觉”不是模型所具备的属性,而是系统所强制执行的属性。提出了HALO(幻觉感知分层监督),一种将幻觉视为可控制的故障模式而非可消除故障模式的保证架构。HALO由六层防御组成,详细介绍了每层,特别关注基于证据的置信度,并在规范的索赔提取工作负载上说明了该架构。

英文摘要

Enterprises will not deploy AI agents they cannot trust, and the most-cited reason for distrust is hallucination: confident, fluent output that is simply not true. The common response is to wait for a model that does not hallucinate. We argue that this is the wrong target. Large language models are, by construction, capable of generating unsupported text, and no amount of scale removes the possibility; a faithfulness judge bolted onto a raw model catches some errors but still ships others, and even well-curated retrieval pipelines have been shown to fabricate citations. We reframe the goal: "zero hallucination" is not a property a model possesses but a property a system enforces. We present HALO (Hallucination-Aware Layered Oversight), an assurance architecture which treats hallucination as a containable failure mode rather than an eliminable one. HALO composes six layers of defense: grounded generation over retrieved, approved content; constrained, deterministic execution that bounds where the model can err; multi-signal verification that scores every output for groundedness and hallucination using both an LLM judge and evidence-based checks against the source text; calibrated abstention, so the system declines rather than guesses when grounding is insufficient; total traceability of every retrieval, tool call, and generation; and continuous oversight that detects drift, alerts on threshold breaches, and closes the loop by regenerating and statistically validating improved agents. We detail each layer, give particular attention to evidence-based confidence (which verifies extractions against the source document rather than trusting the model's self-reported certainty), and illustrate the architecture on a regulated claims-extraction workload.

Comments14 pages, 2 figures

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑