随机语义证据图:智能体AI的不确定性传播与治理
Finding Icebergs in Language-Model Workflow: Diagnosing Latent Structural Fragility with Stochastic Semantic Evidence Graphs
浏览论文内容
中文总结 AI 辅助
提出随机语义证据图(SSEG),通过层次化随机DAG传播不确定性并界定终端误差,用于诊断智能体AI的治理触发点,实验验证了其有效性。
中文摘要 AI 辅助
AI智能体的评估通常只检查最终答案,然而错误可能通过证据、检索、提示、生成或决策映射等环节引入。我们提出了一种随机语义证据图(SSEG),这是一个层次化的随机有向无环图(DAG),其语言节点扩展为自回归令牌子图,其可观测输出可能是完整短语上的一个分布。语义简化和校准是可选的。我们定义了图相关的局部缺陷和下游边影响,推导了终端误差的路径上界,并利用其节点级项来诊断治理触发条件。对于来源溯源,该图保留了不确定的声明-段落关系,并传播尖锐的弗雷歇界,而非假设各来源独立。在三种开放权重架构中,信息等价的变化实质性地改变了完整短语的分布。一项受控实验在5,000个案例中未出现证书违规;交叉RAG和实时Brave检索实验分离了检索、呈现、来源和交互效应。因此,SSEG将工作流溯源转化为对不确定性进入位置、传播方式以及输出是否合格使用的定量描述。
英文摘要
AI-workflow governance cannot be reduced to checking the final answer: an apparently safe answer may rest on a fragile evidence path that ordinary evaluation cannot see, localize or govern. We call this hidden fragility a "structural iceberg": hallucinations and unsupported claims may form its visible tip, while consequential weakness remains submerged. Stochastic semantic evidence graphs (SSEGs) expose these icebergs by preserving workflow channels, propagating local uncertainty and identifying the hidden paths on which an apparently safe output depends. ALCE and RAGTruth show that visible failures at the tip -unsupported citations and hallucinated spans -rest on distinct submerged weaknesses and therefore require different interventions. Across retrieval, tool-use and controlled stress tests, SSEG localizes those weaknesses, supports targeted repair, produces no false automatic passes in 35,000 known-truth cases and reduces ToolSandbox review by 28.8% across 96 executions from two agent models. The same structural view carries into end-to-end governance: in a separately sealed 1,200-case FinGovBench study, adding SSEG to GPT-OSS-20B reduces unsafe releases from 452/660 to 8/660 while releasing all 540 safe cases and correctly distinguishing 592/600 matched workflow pairs. An unchanged-gate transfer to Qwen3-8B releases all 540 safe cases and none of 660 unsafe cases, whereas flat-UQ releases 520 unsafe cases. SSEG therefore moves governance below surface-level output checking, turning hidden evidence dependencies into auditable, path-specific decisions about intervention, revalidation and release.
发表机构
- BeliefLens
机构由 AI 辅助整理,请以论文原文为准。