政策审计报告的白盒证据包
White Box Evidence Packages for Policy Audit Reports
浏览论文内容
中文总结 AI 辅助
研究在政策审计中评审者如何判断报告有无证据支持,引入可控评估框架,在多种证据条件下生成报告让评审员评估。结果显示内部证据影响报告引用推理方式,混合接口最有用,随机控制揭示风险,还将内部模型访问重构为证据设计问题。
中文摘要 AI 辅助
随着人工智能治理从基准分数转向可审计监督,一个核心问题是评审者如何判断大型语言模型生成的审计报告是否有证据支持。本文在基于段落的政策审计中研究该问题,报告须解读给定政策段落并为其主张引用证据。我们引入一个可控评估框架,固定段落、评分标准和审计模型,仅改变提供给审计师的证据接口。在60个AGORA政策案例中,我们在十种证据条件下生成600份结构化报告,包括基于段落的证据、内部模型证据等。五名人类评审员评估主要接口的正确性、段落基础、诊断有用性和证据滥用情况。结果表明内部证据改变报告引用和推理证据的方式,但更多内部引用本身不会使报告更有效。白盒诊断解释了失败模式:因果定位狭窄,而报告容易重复使用更宽泛的可读标签和令牌方向。混合接口平均最有用,而随机控制揭示了一个关键治理风险:报告引用无关内部证据时听起来可能很合理。本研究将内部模型访问重新构建为审计工作流程的证据设计问题,而非透明度保证。
英文摘要
As AI governance moves from benchmark scores toward auditable oversight, a central question is how reviewers can tell whether an LLM-generated audit report is actually supported by evidence. This paper studies that question in passage-anchored policy audits, where a report must interpret a given policy passage and cite evidence for its claims. We introduce a controlled evaluation framework that holds the passage, rubric, and auditor model fixed while changing only the evidence interface supplied to the auditor. Across 60 AGORA policy cases, we generate 600 structured reports under ten evidence conditions, including passage-based evidence, internal model evidence, a hybrid package, and a shuffled control that preserves evidence format while breaking case relevance. Five human reviewers evaluate the primary interfaces for correctness, passage grounding, diagnostic usefulness, and evidence misuse. The results show that internal evidence changes how reports cite and reason about evidence, but more internal citations do not by themselves make a report more valid. A white-box diagnostic explains the failure mode: causal localization is narrow, while reports readily reuse broader readable labels and token directions. The hybrid interface is the most useful on average, while the shuffled control exposes a key governance risk: reports can sound substantively plausible while citing irrelevant internal evidence. This study reframes internal model access as an evidence design problem for audit workflows, rather than as a guarantee of transparency.