发表机构
MIT(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对统计学习算法推断人类偏好的挑战,提出‘从权重到文字’方法,输入选择问题数据集,自动发现偏好维度,解决欠定和不透明问题,经多领域展示及人体实验验证,能提高偏好模型预测准确性,获参与者认可。
AI 中文摘要
统计学习算法从高维选择数据推断人类偏好面临挑战,选择方案多因素并存致难以确定驱动决策的因素,且方法不透明。我们引入‘从权重到文字’方法,以选择问题数据集为输入,自动发现与领域相关的偏好维度,用自然语言描述并与模型表示空间中的向量配对。该方法解决了欠定和不透明问题,可集中归因于少量有意义因素,用自然语言外化模型推理以便用户实时检查和编辑。我们先在四个不同领域定性展示其通用性,然后报告了两个预注册的人体实验,证明其对学习偏好模型的益处:使偏好模型向学习到的基础正则化可提高对保留选择的预测准确性,纳入参与者的结构化编辑可进一步提高准确性。在直接比较中,参与者更喜欢该方法推断的偏好概况,并认可其预测更准确。
英文摘要
The growing use of statistical learning algorithms to infer human preferences from high-dimensional choice data runs up against a fundamental challenge: choice alternatives typically differ in many ways simultaneously, so it is generally unclear which factors actually drove an observed decision and should be credited as preferences. Compounding this problem, the opacity of these methods leaves human operators unable to inspect, contest, or correct models when they err. We introduce \emph{weights to words}, a method that takes a dataset of choice problems as input and automatically discovers a collection of domain-relevant preference dimensions, each described in natural language and paired with a vector in the model's representational space. These dimensions address both under-determination and opacity: they can be applied to concentrate attribution on a small set of meaningful factors, and they can externalize the model's inferences in natural language so that users can inspect and edit them in real time. We first qualitatively illustrate the method's versatility on four diverse domains: moral dilemmas, movies, wines, and free-form LLM responses. We then report two pre-registered human-subjects experiments, on moral dilemmas ($N=450$) and movie selection ($N=449$), that demonstrate its benefits for learning preference models: (1) regularizing a preference model toward the learned basis increases prediction accuracy on held-out choices, and (2) incorporating participants' structured edits further improves accuracy. In head-to-head comparisons, participants prefer the method's inferred preference profiles and endorse its predictions as more accurate.
Comments42 pages, 22 figures, 14 tables; main text 11 pages, remainder appendices