arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

标签一致性不能衡量授权

Label Agreement Does Not Measure Authorization

Amir Sabbaghziarani, Bradley Thomas Baker, Theodore J. LaGrow, Sergey Plis

arXiv 2610.04544首次发表:更新:

发表机构

Georgia State University; Georgia Institute of Technology; Emory University(佐治亚州立大学; 佐治亚理工学院; 埃默里大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

研究发现标签一致性无法衡量智能体LLM流水线的授权、完整性和证据敏感性,提出分离测量并配对对照以暴露失败。

AI 中文摘要

许多团队现在将标签本体和元数据协调工作委托给智能体LLM流水线。我们构建了这样一个流水线并对其进行了审计。我们的总体得分看起来健康,但流水线却以这些得分未显示的方式持续失败,因此我们着手找出它们隐藏了什么。标签一致性询问的是提议的标签是否与参考匹配。它不询问智能体是否有权提议该标签,输出是否足够完整以供执行,或者当证据变化时标签是否随之变化。我们在COBRE和FBIRN上分别测量了这三个方面,这两个数据集是来自不同联盟的精神分裂症和对照组神经影像队列,它们与标签一致性以及彼此之间都存在分歧。向智能体展示上游提议几乎不改变标签一致性,从0.857到0.870,而所选行动的一致性翻倍,从0.409到0.830。解析为JSON的输出仍在一个模型的10%案例和另一个模型的33%案例中丢失必需字段。而且,重放其首次回答的智能体在原始案例上得分完美,但一旦我们改变决定它们的证据,得分即为零。在下游,这些指标均未报告的行序错误抹去了大部分诊断信号。因此,我们分开测量这些属性,将每个属性与对照配对,并根据结果门控承诺,这使得失败可见且易于路由给人工处理。这些都不能防止失败。我们的参考标签是规则派生的,因此与它们的一致性意味着一致性而非正确性。代码可在https URL获取。

英文摘要

Many groups now delegate label ontology and metadata harmonization to agentic LLM pipelines. We built one and audited it. Our aggregate scores looked healthy, but the pipeline kept failing in ways they did not show, so we set out to find what they hid. Label agreement asks whether a proposed label matches a reference. It does not ask whether the agent was entitled to propose it, whether the output was complete enough to act on, or whether the label moved when the evidence moved. We measured those three separately on COBRE and FBIRN, two schizophrenia and control neuroimaging cohorts from different consortia, and they come apart, from label agreement and from each other. Showing the agent an upstream proposal barely moves label agreement, 0.857 to 0.870, while agreement on the chosen action doubles, 0.409 to 0.830. Output that parses as JSON still drops a required field on 10% of one model's cases and 33% of the other's. And an agent that replays its first answer scores perfectly on original cases and zero once we change the evidence that decides them. Downstream, a row-order error that none of these metrics reports erases most of the diagnostic signal. So we measure these properties apart, pair each with a control, and gate commitment on the result, which makes failures visible and easy to route to a person. None of this prevents failure. Our reference labels are rule-derived, so agreement with them means consistency, not correctness. Code is available at https://github.com/amir-sbg/Label-Agreement-Does-Not-Measure-Authorization.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑