arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

审计规则如何塑造LLM中的忠实因子解释

How the Audit Rule Shapes Faithful Factor Explanations in LLMs

Taolin Zhang, Hanyu Wang, Jiuheng Wan, Tingyuan Hu, Chengyu Wang

arXiv 2610.01514首次发表:更新:

发表机构

Hefei University of Technology; East China Normal University; Alibaba Cloud Computing(合肥工业大学; 华东师范大学; 阿里云计算)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过验证博弈框架,揭示有限预算下审计规则对LLM因子解释报告的影响,提出不依赖报告的审计组件可抑制少报行为,并用CBS在四个NLP基准上验证。

AI 中文摘要

大型语言模型经常被询问哪些输入因子影响了它们的输出。对于结构化输入,此类报告可以通过反事实扰动进行核查,但每个因子必须被多次查询以估计其效应,因此验证通常受预算限制。我们研究了这种有限预算设置如何改变如实报告因子级影响力的激励。我们将这种交互形式化为一个验证博弈,并表明当审计依赖于报告时,仅凭适当评分是不够的:依赖于报告的审计会产生抑制激励,因为被报告为重要的因子更可能被检查并因估计噪声而受到惩罚。相反,不依赖于报告的审计,或带有少量不依赖于报告的底线的混合规则,消除了这一渠道,并使如实报告优于完全抑制。我们用反事实布里尔分数(CBS)实例化该框架,并在四个NLP基准上评估其预测。一个合成理性智能体精确匹配理论预测,而真实LLM在激励被明确化时遵循相同的激励。主要的设计启示很简单:在部分验证下,因子级解释系统应包含一个不依赖于报告的审计组件,以便少报不能用来避免审查。

英文摘要

Large language models are often asked which input factors influenced their outputs. For structured inputs, such reports can be checked by counterfactual perturbation, but each factor must be queried multiple times to estimate its effect, so verification is usually budget-limited. We study how this limited-budget setting changes the incentive to report factor-level influence truthfully. We formalize the interaction as a verification game and show that proper scoring alone is not enough when auditing depends on the report: report-dependent auditing creates a suppression incentive, because factors reported as important are more likely to be checked and penalized for estimation noise. In contrast, report-independent auditing, or a mixed rule with a small report-independent floor, removes this channel and makes truthful reporting preferable to full suppression. We instantiate the framework with the Counterfactual Brier Score (CBS) and evaluate its predictions on four NLP benchmarks. A synthetic rational agent matches the theoretical prediction exactly, and real LLMs follow the same incentives when they are made explicit. The main design implication is simple: under partial verification, factor-level explanation systems should include a report-independent audit component so that under-reporting cannot be used to avoid scrutiny.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑