arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

LLM 法官评估实体对齐的可靠性

Reliability of LLM Judges for Evaluating Entity Alignment

Vaibhava Lakshmi Ravideshik, Mayank Kejriwal

arXiv 2610.09554首次发表:更新:

发表机构

University of Michigan, Ann Arbor; Information Sciences Institute, University of Southern California(密歇根大学安娜堡分校; 南加州大学信息科学研究所)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究首次系统评估 LLM 法官在实体对齐任务中的可靠性,发现标签暴露导致锚定偏差削弱判别能力,提出无标签协议可恢复性能,并发布生物医学基准与审计框架。

AI 中文摘要

实体对齐(EA)识别跨知识图谱的等价实体,对于知识库集成和本体合并至关重要。大规模评估 EA 系统需要昂贵的专家标注,使得跨不同领域的系统性评估在实践上不可行。LLM 作为法官的评估提供了一种可能可扩展的替代方案,但其在 EA 等结构化预测任务中的可靠性仍未得到研究。我们提出了首个系统性基准研究,涵盖三个前沿模型、三个数据集和四个 EA 系统,使用扰动偏差诊断、跨所有数据集-法官-提示组合的元评估以及反事实标签翻转测试。我们识别出锚定偏差,即当系统决策标签可见时,法官反转判别的一种失败模式。标签暴露因果性地削弱法官判别能力(J-ROC-AUC 0.12-0.87),而无标签协议在独特名称数据集上恢复接近上限的能力(0.93-1.00),在生物医学对上显著恢复(0.93-0.95)。反事实实验确认了因果性(FSR 53-99%),并揭示了一个前沿模型悖论:更强的法官表现出更大的标签敏感性,而非更小。一项双盲双标注者人工评估(102 对,Cohen's kappa=0.902)直接确认了这一机制。我们发布了首个生物医学 EA 基准(MeSH-SNOMED CT,15K 对)和一个可复现的 LLM 法官可靠性审计框架。代码和数据可在该 https URL 获取。

英文摘要

Entity Alignment (EA) identifies equivalent entities across knowledge graphs and is critical for knowledge base integration and ontology merging. Evaluating EA systems at scale requires expensive expert annotation, making systematic assessment across diverse domains practically infeasible. LLM-as-judge evaluation offers a potentially scalable alternative, yet its reliability for structured prediction tasks like EA remains unstudied. We present the first systematic benchmarking study across three frontier models, three datasets, and four EA systems, using perturbation bias diagnostics, meta-evaluation across all dataset-judge-prompt combinations, and counterfactual label-flip tests. We identify anchor bias, a failure mode in which judges invert discrimination when the system's decision label is visible. Label exposure causally collapses judge discrimination (J-ROC-AUC 0.12-0.87), while a label-free protocol recovers near-ceiling capability on distinctive-name datasets (0.93-1.00) and significant recovery on biomedical pairs (0.93-0.95). Counterfactual experiments confirm causality (FSR 53-99%) and reveal a frontier model paradox: stronger judges exhibit greater label sensitivity, not less. A blinded two-annotator human evaluation (102 pairs, Cohen's kappa=0.902) confirms this mechanism directly. We release the first biomedical EA benchmark (MeSH-SNOMED CT, 15K pairs) and a reproducible auditing framework for LLM judge reliability in EA. Code and data are available at https://github.com/vaibhavalakshmiravideshik/llm-as-a-judge-entity-alignment.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑