超越准确性:程序性痕迹如何改变LLM监督者的决策标准
Beyond Accuracy: How Procedural Traces Shift the Decision Criterion of LLM Overseers
浏览论文内容
中文总结 AI 辅助
本研究通过信号检测理论审计LLM监督者,发现程序性痕迹细节会改变决策标准而非准确性,增加误报,建议AI审计员应评估决策标准和误报行为。
中文摘要 AI 辅助
组织越来越多地使用监督循环,其中一个大型语言模型(LLM)在声称步骤的程序性痕迹伴随下审计另一个模型的输出。对此类LLM作为评判者流程的一个常见担忧是,详细的痕迹会使监督者变得轻信。利用信号检测理论,我们在19项合规任务(4,551个分析判断)上审计了五个LLM监督者,仅改变痕迹细节和证据标签。当否定性证据始终可见时,错误检测保持在接近上限的水平。相反,精细的痕迹将决策标准转向拒绝,增加了易受影响监督者的误报。在没有选项标签的情况下,经人工验证的理由编码显示约60%的误报引用了无法将证据与其选项关联的问题。标签消除了这一陈述的理由,但这些监督者对正确工作的残余拒绝仍然存在,并随痕迹细节增加而上升。因此,程序性痕迹作为治理工件,塑造了监督决策。AI审计员应通过其决策标准和误报行为以及准确性来评估。
英文摘要
Organizations increasingly use oversight loops where one large language model (LLM) audits another's outputs alongside procedural traces of claimed steps. A common concern about such LLM-as-a-judge pipelines is that detailed traces make overseers gullible. Using signal detection theory, we audit five LLM overseers on 19 compliance tasks (4,551 analyzed judgments), varying only trace detail and evidence labeling. With disconfirming evidence always visible, error detection remains near ceiling. Instead, elaborate traces shift the decision criterion toward rejection, increasing false alarms in susceptible overseers. Without option labels, human-validated reason coding shows about 60% of false alarms cite an inability to tie evidence to its option. Labels eliminate this stated reason, yet residual rejection of correct work persists in those overseers and rises with trace detail. Procedural traces thus act as governance artifacts that shape oversight decisions. AI auditors should be evaluated by their decision criterion and false-alarm behavior, alongside accuracy.
发表机构
- Stevens Institute of Technology(史蒂文斯理工学院)
- University of Massachusetts Boston(马萨诸塞大学波士顿分校)
- Stony Brook University(石溪大学)
机构由 AI 辅助整理,请以论文原文为准。