发表机构
Carnegie Mellon University; Microsoft(卡内基梅隆大学; 微软)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究通过监督微调训练语言模型保持信念并最大化预期效用,在20个数据集上验证其跨领域和框架的决策泛化能力,并引入信息不完整设置测试缺失信息识别,证明针对性微调能显著提升连贯决策并支持可决策性判断。
AI 中文摘要
可靠的决策需要的不只是准确的预测:模型必须保持其信念,应用相关的效用,并识别何时缺少证明行动所需的信息。我们研究语言模型是否能通过监督微调学习这一决策过程,并将其泛化到不同领域和决策挑战的不同自然语言表达中。在20个数据集上,我们通过首先引出结果概率,然后在保持证据不变的情况下仅改变效用和决策问题的框架,来探索信念不稳定和决策错误的挑战。我们训练模型在保持引出的信念的同时选择最大化预期效用的行动,并评估向未见过的应用领域、保留的框架和不同类别的支付结构的迁移。我们进一步引入了信息不完整的设置,其中所需的效用被扣留并替换为无关文本,测试模型能否区分缺失的决策相关信息与仅仅是额外的上下文。我们发现,有针对性的微调显著改善了连贯的决策,并且在许多情况下,学习会跨领域和框架迁移到训练期间未观察到的情况。此外,为可决策性训练的模型学会识别何时无法基于缺失信息证明行动的合理性。最后,我们展示了路由系统的价值,该系统分别考虑决策完整性的识别和效用敏感的决策执行。
英文摘要
Reliable decision-making requires more than accurate prediction: a model must preserve its beliefs, apply the relevant utilities, and recognize when the information needed to justify an action is missing. We study whether language models can learn this decision procedure from supervised fine-tuning and generalize it across domains and differing natural-language expressions of the decision challenge. Across 20 datasets, we explore challenges of belief instability and decision-making errors by first eliciting probabilities of outcomes and then varying only the utilities and the framing of the decision problems, while holding the evidence fixed. We train models to preserve elicited beliefs while selecting the action that maximizes expected utility, and evaluate transfer to unseen application domains, held-out framings, and different classes of payoff structures. We further introduce incomplete-information settings in which required utilities are withheld and replaced with irrelevant text, testing whether models can distinguish missing decision-relevant information from merely additional context. We find that targeted fine-tuning substantially improves coherent decision-making and that in many situations, learning transfers across domains and framings to situations unobserved during training. Further, models trained for decidability learn to identify when action cannot be justified based on missing information. Finally, we show the value of a routed system that considers separately the recognition of decision completeness and utility-sensitive decision execution.