arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05015econ.THcs.AIcs.LG

显示理性:基于表示定理的无标签评估与正则化

Revealed Rationality: Label-Free Evaluation and Regularization from Representation Theorems

Isaiah Andrews

首次发表
浏览论文内容

中文总结 AI 辅助

本文利用决策理论表示定理的“当且仅当”结构,提出对LLM等AI系统的无标签评估与正则化方法,通过三个实例生成零值连续惩罚项,补充其他训练信号。

中文摘要 AI 辅助

决策理论中的表示定理确立,行为满足特定公理当且仅当它可由明确定义的目标合理化。本文认为这种“当且仅当”结构为大型语言模型(LLM)及其他AI系统的无标签评估与正则化提供了潜在有用的基础。公理符合性可通过模型自身对合成选择问题的响应来检查,无需外部标签或人类反馈,且惩罚项易于计算。由于公理是必要且充分的,所得检查穷尽了相关理性标准对所引出数据的所有含义:通过检查的模型,无法基于理性理由被同一数据的任何进一步测试否定。本文讨论了三个实例:基于德菲内蒂(de Finetti)定理的概率一致性、基于阿夫里亚特(Afriat)定理的偏好理性,以及基于埃切尼克和斋藤(Echenique and Saito,2015)定理的主观期望效用,每个实例均产生一个连续惩罚项,当行为可被合理化时该惩罚项为零。由于一致性不限制使行为合理化的目标,这些惩罚项补充而非替代其他评估与训练信号。

英文摘要

Representation theorems in decision theory establish that behavior satisfies certain axioms if and only if it can be rationalized by a well-defined objective. I argue that this ``if and only if'' structure provides a potentially useful foundation for label-free evaluation and regularization of LLMs and other AI systems. Axiom compliance can be checked from the model's own responses to synthetic choice problems, with no external labels or human feedback, and the penalties are readily computable. Because the axioms are necessary and sufficient, the resulting checks exhaust the implications of the relevant rationality standard for the elicited data: a model that passes cannot be rejected on rationality grounds by any further test of the same data. I discuss three instantiations: probabilistic coherence via a theorem of de Finetti, preference rationality via Afriat's theorem, and subjective expected utility via a theorem of Echenique and Saito (2015), each yielding a continuous penalty that is zero whenever behavior can be rationalized. Since coherence does not restrict which objective rationalizes behavior, these penalties complement rather than replace other evaluation and training signals.

发表机构

  • Massachusetts Institute of Technology(麻省理工学院)
  • NBER(美国国家经济研究局)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑