大语言模型在不确定性下的偏好推理
Preference Reasoning under Indeterminacy in Large Language Models
- Penn State University(宾夕法尼亚州立大学)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文聚焦大语言模型决策智能体的偏好推理问题,将不确定性挑战形式化为认知与结构两类,发现当前模型无法区分确定与不确定实例,推理校准不当。
AI中文摘要:
随着大语言模型发展为决策智能体,偏好推理能力对于对齐、协作和集体智能而言至关重要。但与标准基准不同,现实世界中的偏好推理天生具有不确定性:信息可能不完整,且可能不存在有效解决方案。本文认为,不确定性而非仅正确性,是AI推理的核心挑战。我们从两个维度将该挑战形式化:(i)认知不确定性,源于不完整、部分或表达性偏好;(ii)结构不确定性,源于标准社会选择概念下不存在解决方案。在一系列任务中,我们发现最先进的语言模型会系统性地无法区分确定与不确定实例,即便在验证场景中也表现出校准不当的推理。
英文摘要:
As large language models evolve into decision-making agents, the ability to reason over preferences becomes fundamental to alignment, coordination, and collective intelligence. Yet, unlike standard benchmarks, real-world preference reasoning is inherently indeterminate: information may be incomplete, and valid solutions may not exist. We argue that indeterminacy, rather than correctness alone, is a central challenge for AI reasoning. We formalize this challenge along two axes, (i) epistemic indeterminacy, arising from incomplete, partial, or expressive preferences, and (ii) structural indeterminacy, arising from the non-existence of solutions under standard social choice concepts. Across a hierarchy of tasks, we show that state-of-the-art language models systematically fail to distinguish between determined and undetermined instances, exhibiting miscalibrated reasoning even in verification settings.