发表机构
IRIT, Université de Toulouse; IPAL-CNRS-NUS-A*STAR; IRIT, Université Toulouse Capitole(图卢兹大学图卢兹计算机研究所; IPAL-CNRS-NUS-A*STAR; 图卢兹 Capitole 大学图卢兹计算机研究所)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究利用层次分析法构建基准并提出端到端方法,使大型语言模型能透明地执行完整多准则决策流程,在法律和高等教育排名中显著提升与专家判断的一致性。
AI 中文摘要
大型语言模型(LLMs)越来越多地被应用于广泛的决策任务中。然而,其内部推理的不透明性使得验证或解释其输出变得困难,而在高风险场景中,可解释性的需求变得尤为关键。本研究通过层次分析法(AHP)——一种经典且广泛使用的多准则决策框架——来考察LLMs的决策能力。我们基于AHP构建了一个新的带标注基准,并提出了首个端到端方法,使LLMs能够执行完整的AHP工作流程。在法律和高等教育排名领域的真实世界决策问题实验中,我们的方法显著提高了与专家判断的一致性。
英文摘要
LLMs are increasingly employed in a wide range of decision-making tasks. However, the opacity of their internal reasoning makes it difficult to validate or interpret their outputs, and the need for interpretability becomes especially critical in high-stakes settings. This study examines the decision-making capabilities of LLMs through the Analytic Hierarchy Process (AHP), a classical and widely used multicriteria decision-making framework. We construct a new annotated benchmark based on AHP and propose the first end-to-end approach that enables LLMs to perform the complete AHP workflow. Experiments in real-world decision problems in the legal and higher-education ranking domains show that our method significantly improves alignment with expert judgments.