大语言模型(LLM)得出的偏好判断不具备自一致性
LLM-Derived Preference Judgments Are Not Self-Consistent
浏览论文内容
中文总结 AI 辅助
研究发现,用于解读人类偏好的LLM得出的数值化偏好判断存在大量持续不一致性,无法由单一效用函数概括,打破了相关研究的自一致性假设。
中文摘要 AI 辅助
智能体越来越多地通过查询大语言模型(LLM)获取数值化偏好判断来解读人类的自然语言偏好,例如询问人类愿意为某商品支付的金额。越来越多的研究从这些判断中估计出效用函数,再基于估计的效用选择行动。该流程假设这些判断大致具备自一致性,即存在单一效用函数可复现这些判断。但事实是否如此?为研究该问题,我们对基数型LLM偏好判断的自一致性进行了测量,例如,两件商品的 stated 支付意愿差异应等于使人类对交换这两件商品无差异的 stated 支付金额。我们开发了统计检验方法和可解释指标,用于衡量观测到的响应与最佳拟合自一致效用函数的偏离程度。针对六个LLM的航班、公寓、酒店示例实验显示存在大量持续的不一致性。这表明,LLM得出的偏好判断无法由单一效用函数忠实地概括。
英文摘要
Agents increasingly interpret a person's natural-language preferences by querying an LLM for numerical preference judgments, e.g., by asking how much the person would be willing to pay for an item. A growing body of work estimates a utility function from these judgments and then chooses actions based on their estimated utility. This pipeline assumes the judgments are approximately self-consistent: that a single utility function can reproduce them. But are they? To study this question, we measure the self-consistency of cardinal LLM preference judgments. For example, the difference in stated willingness-to-pay between two items should match the stated payment that makes a person indifferent to exchanging them. We develop statistical tests and interpretable measures of how far observed responses depart from the best-fitting self-consistent utility function. Experiments with flight, apartment, and hotel examples across six LLMs reveal large persistent inconsistencies. This suggests that LLM-derived preference judgments cannot be faithfully summarized by a single utility function.