arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.14329cs.CRcs.AIcs.CLcs.CYcs.LG

面向基于原则的监管的LLM作为评判者的四轴可信度基准

A Four-Axis Trustworthiness Benchmark for LLM-as-Judge in Principle-Based Regulation

发表机构Independent Researcher
查看机构详情
  • Independent Researcher

机构由 AI 辅助整理,请以论文原文为准。

Dipankar Sarkar

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出面向基于原则监管的LLM评判者的四轴可信度基准,发布Principle-Bench并引入Ceca评估器,发现无方法在所有四轴占优,部署级LLM评判者需报告对抗性欺骗与事后校准等指标。

中文摘要 AI 辅助

基于原则的监管(如“公平、清晰且不具误导性”或“产生良好结果”等评估标准)无法简化为二元谓词,因此大型语言模型(LLM)作为评判者正日益被用作替代方案。本文提出,此类评判者必须在四个轴上接受评估:准确性、释义鲁棒性、对抗鲁棒性和校准。我们发布了Principle-Bench,该基准包含168个加密资产金融推广场景,映射到英国金融行为监管局(FCA)的两项原则,且包含根据预注册规则生成的释义、对抗性关键词堆砌和边界扰动,是首个覆盖基于原则监管全部四个轴的基准。我们还引入了Ceca(校准示例-集群评估):一种可校准、可审计的评估器,能生成精确的每个示例的反事实归因。在关键词计数、三个句子嵌入器、一个开放权重LLM评判者和一个校准级联中,没有任何方法在所有四个轴上都占优。一个120B的LLM评判者在良性输入上表现最强,但在关键词堆砌的《消费者责任》输入上准确性下降了47个点(从0.74降至0.27),表现为“合规假象”。来自不同模型系列的第二个评判者在该拆分上仅以Cohen’s kappa=0.16达成一致,将失败归因于模型而非语料库。任何用于基于原则监管的部署级LLM评判者必须报告每个原则的对抗性欺骗和事后校准,以及总体准确性。

英文摘要

Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute. Our position is that any such judge must be evaluated on four axes: accuracy, paraphrase robustness, adversarial robustness, and calibration. We release Principle-Bench, 168 cryptoasset financial-promotion scenarios mapped to two UK FCA principles, with paraphrase, adversarial keyword-stuffing, and boundary perturbations authored under a pre-registered rubric; the first benchmark covering all four axes for principle-based regulation. We also introduce Ceca (Calibrated Exemplar-Cluster Assessment): a calibrated, auditable assessor that emits exact per-exemplar counterfactual attributions. Across keyword counting, three sentence-transformer embedders, an open-weight LLM-judge, and a calibrated cascade, no method dominates all four axes. A 120B LLM-judge, strongest on benign inputs, loses 47 accuracy points (0.74 to 0.27) on keyword-stuffed Consumer Duty inputs: "compliance theatre." A second judge from a different model family agrees only at Cohen's kappa = 0.16 on that split, localising the failure to the model rather than the corpus. Any deployment-grade LLM-judge for principle-based regulation must report per-principle adversarial deception and post-hoc calibration alongside aggregate accuracy.

补充信息

↑