立场:是时候针对自一致性优化大型语言模型(LLMs)了
Position: It's Time to Optimize LLMs for Self-Consistency
浏览论文内容
中文总结 AI 辅助
本文提出自一致性框架,指出LM的缺陷源于单输出对独立评估的假设,可通过一致性优化解决,为开发通用一致性LLM提供思路。
中文摘要 AI 辅助
尽管语言模型(LM)的预训练与后训练流程日益复杂,仍存在诸多重要缺陷:模型过度依赖用户表述(“奉承性”)、逻辑泛化不完整,以及生成自信但错误的响应。本文认为,这些缺陷源于贯穿整个流程的建模假设:即行为可基于单输出对独立指定与评估。若不考量模型在不同输入下的响应关系,许多模型缺陷即便并非无法检测,也极难被发现。在这篇立场论文中,我们提出自一致性作为理解这些缺陷的框架。我们首先观察到,各类旨在改进LM行为特定方面的技术——涵盖对抗鲁棒性、事实连贯性等不同属性——均可被理解为通用“一致性优化”流程的特例,且可通过一套标准优化工具解决。随后,我们概述了通过优化一致性可实现的一系列新模型属性,并在最后探讨开发通用一致性LLM的意义,包括其将具备的能力及引发的质疑。
英文摘要
Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but incorrect responses. We argue that these failures arise from a modeling assumption permeating all aspects of the pipeline: that behavior can be specified and evaluated independently on single-output pairs. Many model failures are difficult, if not impossible, to detect without reasoning about relationships between a model's responses across inputs. In this position paper, we propose self-consistency as a framework for understanding these failures. We first observe that a wide variety of techniques designed to improve specific aspects of LM behavior-targeting properties as diverse as adversarial robustness and factual coherence-can be understood as special cases of a common "consistency optimization" procedure and addressed with a standard set of optimization tools. We next outline a set of new model properties that could be achieved by optimizing for consistency, and conclude with a discussion of what it would mean to develop generally consistent LMs, including the capabilities they would enable and the objections they raise.