arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.13926cs.AIcs.CLcs.DB

绝非数字:将答案作为事实使用的AI系统的结构弃权

Never the Number: Structural Abstention for AI Systems Whose Answers Are Consumed as Fact

  • Apple Inc.(苹果公司)

机构由 AI 辅助整理,请以论文原文为准。

Zhelun, Wu

中文总结 AI 辅助

针对LLM文本转SQL系统返回错误答案且无法区分的问题,提出带生成式外壳的可信内核架构,即结构弃权,在生产案例中验证其可靠性优于两种生成式替代方案。

中文摘要 AI 辅助

大型语言模型使数据库自然语言接口(NLIDB)重新获得了可信度,但LLM文本转SQL系统在部署时存在关键问题:幻觉生成的列或聚合错误的总数会产生流畅的错误答案,在使用时与正确答案无法区分。当使用者无法检查生成的查询时,如企业AI部署和运营仪表板,且使用者越来越多地是使用工具的智能体而非人类,仅靠准确性是不够的:没有任何标记可用于判断哪些答案值得不信任。这首先是一个可靠性问题,其次才是准确性问题。我们为此类系统提出了一种架构模式:带有生成式外壳的可信内核,基于一个不变量:能够生成内容的组件可影响系统回答哪个问题,但绝不会影响其返回哪个值。生成式外壳解释未明确的输入并表述回复;确定性内核将明确的问题与可回答的问题形状的有界集合匹配,并通过确定性执行将其编译为查询。两者在用户计算任何值之前读取的确认处交汇,内核无法表达的请求会被拒绝而非近似处理。我们将此称为结构弃权,并将其与选择性预测和校准置信度的统计弃权区分开来:此处的拒绝无需置信度估计,因为无法回答的请求是不可表示的。我们独立于实现指定该模式,给出一个五步决策方案并在三个领域应用,将不变量从返回值扩展到智能体系统的动作,报告了一项为期两年的生产案例研究,以及两种生成式替代方案:微调解析器和工具检索智能体。最后我们针对企业和自那以后发布的可靠性基准进行了讨论。

英文摘要

Large language models have made natural language interfaces to databases (NLIDB) newly credible, but LLM text-to-SQL systems fail in a way that matters for deployment: a hallucinated column or a mis-aggregated total yields a fluent wrong answer, indistinguishable at the point of use from a right one. Where the consumer cannot inspect the generated query, as in enterprise AI deployments and operational dashboards, and increasingly where the consumer is a tool-using agent rather than a person, accuracy alone is insufficient: nothing marks which answers to distrust. This is a reliability problem before it is an accuracy problem. We propose an architectural pattern for such systems, a trusted kernel with a generative shell, resting on one invariant: a component that can fabricate may influence which question the system answers, never which value it returns. A generative shell interprets underspecified input and phrases replies; a deterministic kernel matches fully specified questions against a bounded set of answerable question shapes and compiles them to queries by deterministic execution. The two meet at a confirmation the user reads before any value is computed, and requests the kernel cannot express are declined rather than approximated. We call this structural abstention, and distinguish it from the statistical abstention of selective prediction and calibrated confidence: refusal here needs no confidence estimate, because unanswerable requests are unrepresentable. We specify the pattern implementation-independently, give a five-decision recipe and work it across three domains, extend the invariant from returned values to the actions of agentic systems, and report a two-year production case study alongside two generative alternatives, a fine-tuned parser and a tool-retrieval agent. We close against enterprise and reliability benchmarks published since.

补充信息

↑