arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

DA-RAC:面向可信AI审计的大语言模型评判器的距离感知校准

DA-RAC: Distance-Aware Calibration of LLM Judges for Trustworthy AI Auditing

Cheng Wu, Vishal Anand, Jaya Krishna Mandivarapu, Xiya Liu, Rui Zhuang

arXiv 2608.14950首次发表:更新:

发表机构

Microsoft(微软公司)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对LLM评判器因无关参考示例导致的校准偏差问题,提出DA-RAC距离感知参考锚定校准方法,在多轮评估基准上提升了校准效果并降低了误通过风险,为可信AI审计提供支撑。

AI 中文摘要

生成式AI系统日益产出现实世界中的产物,然而其效能与有效性常通过无上下文的大语言模型(LLM)评分进行评估。这类评判器可能因无关的上下文参考示例而出现校准偏差,产生错误的置信度,致使低质量或有害输出通过评估。我们将这种失效模式称为上下文诱导的校准偏差,并提出DA-RAC——一种面向LLM评判器的距离感知参考锚定校准方法。DA-RAC为每个评判场景检索语义和结构相似的带标签锚点,按距离对其加权,并将邻域难度作为校准与分流信号。在多轮LLM评判评估基准上,与零样本、思维链评估及静态锚点基线相比,它提升了校准效果并降低了误通过风险。机制分析显示,评判器评分随锚点距离呈系统性变化,而静态参考可能引发误导性决策边界。因此,LLM评判不仅需要更优的模型,还需要校准的、可审计的参考选择,尤其当自动评估用于支持高影响力AI生成产物时;评判应基于相关、可检查且可质疑的解释性产物。

英文摘要

Generative AI systems are increasingly producing real-world artifacts, however their efficacy and validity are often evaluated via context-free LLM-scoring. These judges can be miscalibrated by irrelevant in-context reference examples, creating false confidence and allowing low-quality or harmful outputs to pass evaluation. We study this failure mode as context-induced miscalibration and introduce DA-RAC, a distance-aware reference-anchored calibration method for LLM judges. DA-RAC retrieves semantically and structurally similar labeled anchors for each judgement scenario, weights them by distance, and exposes neighborhood difficulty as a calibration and triage signal. On multi-run LLM-judge evaluation benchmarks, it improves calibration and reduces false-pass risk relative to zero-shot, chain-of-thought evaluation, and static-anchor baselines. Mechanistic analysis shows that judge scores vary systematically with anchor distance, while static references can induce misleading decision boundaries. Thus LLM-judgement requires not only better models, but also calibrated, auditable reference selection, especially when automated evaluation is used to support high-impact AI generated artifacts. Judgments should be grounded in relevant, inspectable, and contestable interpretive artifacts.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑