面向大语言模型的人权基准测试:一种试点方法
Toward Human Rights Benchmarking for LLMs: A Pilot Methodology
AI总结:
该研究开发了首个经专家验证的人权基准HumRightsBench,调整IRAC框架为IRAP,通过真实场景测试LLMs的人权推理能力,发现模型准确率差异显著,可推动AI评估科学发展。
AI中文摘要:
大语言模型(LLMs)日益成为决定哪些人权得以实现、如何实现的中介。然而,目前尚无评估基准可用于衡量LLMs是否能正确推理人权法。为此,我们报告了开发稳健且可扩展方法以创建HumRightsBench的工作:这是首个经专家验证、基于场景的基准,用于评估基于国际人权法义务结构的推理。我们将法律推理的IRAC框架调整为更适配人权工作的独特推理模式(将C“法律结论”替换为P“提出救济措施”,形成IRAP),以构建评估启发式方法。我们还生成了一系列真实场景试点案例,旨在涉及现实世界人权问题的多个维度,并由全球人权律师和专业人员标注。最终,我们发现模型在法律推理任务中的准确率分数差异显著(整体模型性能范围为0.339至0.577,任务的最小-最大范围为0.025至0.774),这有力表明HumRightsBench是在AI评估科学这一新兴子领域发展的关键时期推动其进步的有效工具。
英文摘要:
Large language models (LLMs) increasingly mediate legal determinations over what human rights are realized, and how. Yet, no evaluation benchmark exists to assess whether they can reason correctly about human rights law. To this end, we report our efforts to develop a robust and scalable methodology for creating HumRightsBench: the first expert-validated, scenario-based benchmark for evaluating reasoning grounded in the obligation structure of international human rights law. We adapt the IRAC framework for legal reasoning to better suit the unique reasoning patterns of human rights work (substituting P, "proposing remedies," for C, "legal conclusion," yielding IRAP) to structure our evaluation heuristics. We also produce a pilot series of authentic scenarios designed to implicate the many dimensions of real-world human rights issues and annotated by human rights lawyers and professionals across the world. Ultimately, we find that model accuracy scores range considerably across legal reasoning tasks (overall model performance ranges from 0.339 to 0.577, task min-max ranges from 0.025 to 0.774), which strongly implies that HumRightsBench is a capable instrument for advancing this emerging subfield of AI evaluations science at a critical moment in its evolution.