arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越准确率:面向法律依据任务的视觉语言模型双裁判评估协议

Beyond Accuracy: A Dual-Judge Evaluation Protocol for Vision-Language Models in Legally Grounded Tasks

Su Myat Noe, Ha Thanh Nguyen, May Myo Zin, Ken Satoh

arXiv 2608.24258首次发表:更新:

发表机构

National Institute of Informatics (NII); VinUniversity; ROIS-DS Center for Juris-Informatics(国立信息学研究所; 文大学; ROIS-DS法学信息学中心)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对法律依据任务,提出双裁判评估协议,结合质量与语义等价裁判,在英国交通标志解读任务中发现不对称II型模式,发布相关资源以评估视觉语言模型的法律相关性能。

AI 中文摘要

AI系统在需承担法律责任的场景中接受评估,其正确输出还需符合适用法律标准的论证。现有的法律AI基准和大语言模型(LLM)作为裁判的协议,为衡量任务性能和开放式响应质量提供了重要基础。本文贡献了一种额外的评估信号:双裁判协议,将标准0-10分质量裁判与针对人工整理参考的严格二元语义等价裁判配对。我们研究了一项受控的视觉依据监管任务——英国交通标志解读,其含义是针对每个输入都有已知参考的成文问题,不仅衡量两位裁判的分歧(按设计必然存在),还衡量分歧的程度与位置。在7个可见度等级和2种遮挡模式下的4680次评估中,两位裁判呈中等程度关联(点-双列相关系数r=0.644),同时揭示了影响8.0%所有评估的不对称II型模式。该模式的分布具有启发性:边际比率在高可见度时达到峰值(v=0.8时为14.2%),原因是高分答案在该场景中较为常见;但在答案得分已高于7的条件下,该比率在重度遮挡时最高(v≤0.3时为54-63%),因此高质量得分在输入最退化时最不可信。我们明确该信号是此裁判和参考的属性:一项49行人工检查显示,0-10分裁判与普通读者判断高度一致(皮尔逊相关系数r=0.81;与LLM准确率子得分的r=0.80),而等价裁判则相当但单向地更严格。该协议每次评估增加一次LLM调用,可呈现单裁判协议未报告的信号。我们发布了提示模板、遮挡变体和完整评估结果。

英文摘要

AI systems are increasingly evaluated for legally accountable settings, where correct outputs must also be justifiable against an applicable legal standard. Existing legal-AI benchmarks and LLM-as-judge protocols provide important infrastructure for measuring task performance and open-ended response quality. We contribute one additional evaluation signal: a dual-judge protocol that pairs a standard 0-10 quality judge with a strict binary semantic-equivalence judge against a human-curated reference. We study a controlled, visually grounded regulatory task - UK traffic-sign interpretation, whose meaning is a codified question with a known reference for every input - and measure not merely whether the two judges disagree (by construction they must) but how much and where. On 4,680 evaluations under seven visibility levels and two occlusion modes, the two judges are moderately associated (point-biserial r = 0.644), while revealing an asymmetric Type II pattern affecting 8.0% of all evaluations. Its distribution is instructive: the marginal rate peaks at high visibility (14.2% at v = 0.8) simply because high-scoring answers are common there, but conditioned on the answer already scoring above 7, the rate is highest under heavy occlusion (54-63% at v <= 0.3), so a high quality score is least trustworthy when the input is most degraded. We are explicit that the signal is a property of this judge and reference: a 49-row human check shows the 0-10 judge aligns closely with everyday-reader judgement (Pearson r = 0.81; r = 0.80 with the LLM accuracy sub-score), while the equivalence judge is fairly but one-directionally stricter. The protocol adds one LLM call per evaluation and surfaces a signal single-judge protocols do not report. We release the prompt template, occluded variants, and full evaluation results.

CommentsAccepted and presented at the AI for Law Workshop at ICML 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑