对主观预期效用最大化的敏感性:一项方法学研究,并应用于大语言模型决策
Sensitivity to Subjective Expected Utility Maximization: A Methodological Study, with an Illustrative Application to LLM Decision-Making
浏览论文内容
中文总结 AI 辅助
研究在标记结果稀缺等情况下评估不确定性决策的难题,通过含敏感性参数α的softmax选择模型得出相关可识别性结果,应用于大语言模型决策,检测到结构化比较α效应。
中文摘要 AI 辅助
在标记结果稀缺、成本高昂或与运气混淆的情况下,评估不确定性下的决策很困难。我们将主观预期效用(SEU)最大化视为既定标准,并定义了一种分级度量——SEU敏感性,用于衡量主体对其的符合程度。通过对基于SEU值的备选方案具有敏感性参数α的softmax选择模型进行研究,得出了α以及信念和效用参数(β,δ)的一系列可识别性结果,并通过先验预测检查、参数恢复和基于模拟的校准(SBC)在Stan中进行了验证,同时保留了有限样本的注意事项。在仅包含不确定选择的模型m0中,给定预期效用向量η时α是可识别的且能快速恢复,而(β,δ)的信息则很弱;在扩展模型m1中,原则上δ可通过无β的风险块识别,但在实际样本量下其实际恢复增益可忽略不计,且该块在匹配选择计数时未检测到α精度增益。通过对GPT-4o和Claude 3.5 Sonnet在保险理赔分类和埃尔斯伯格风格瓮问题上的应用,在四个单元格中的两个中检测到了结构化的比较α效应。
英文摘要
Evaluating decisions made under uncertainty is hard when labeled outcomes are scarce, costly, or confounded with luck. We treat subjective expected utility (SEU) maximization as a stated standard and define a graded measure -- SEU sensitivity -- of an agent's conformity to it. The vehicle is a softmax choice model with a sensitivity parameter $α$ on SEU-valued alternatives; the contribution is a sequence of identifiability results for $α$ and for belief and utility parameters $(β, δ)$, validated in Stan via prior predictive checks, parameter recovery, and simulation-based calibration (SBC), with finite-sample caveats intact. In the uncertain-choice-only model $m_0$, $α$ is identifiable given the expected-utility vector $η$ and sharply recovered, while $(β, δ)$ are only weakly informed: the posterior barely contracts and concentrates on a $β$-$δ$ trade-off. In the extended model $m_1$, $δ$ becomes identifiable in principle via a $β$-free risky block, but its practical recovery gain at realistic sample sizes is negligible (matched-count CI-width reduction under 1%), and that block yields no detected $α$-precision gain at matched choice count. These are two distinct phenomena: for $δ$, identifiability does not imply precise estimability at realistic $n$; for $α$, identifiability is silent about what governs finite-$n$ precision. Marginal SBC passes for both models even where the joint posterior is weakly informed -- a demarcation we make precise. A two-by-two application (GPT-4o and Claude 3.5 Sonnet, each on insurance-claims triage and Ellsberg-style urns, with sampling temperature as the lever) runs end-to-end on real LLM choice data, detecting a structured comparative $α$ effect in two of four cells.