arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.05188cs.CLcs.AI

立场:是时候针对自一致性优化大型语言模型(LLMs)了

Position: It's Time to Optimize LLMs for Self-Consistency

Itamar Pres, Belinda Z. Li, Laura Ruis, Zifan Carl Guo, Keya Hu, Mehul Damani, Isha Puri, Ekdeep Singh Lubana, Jacob Andreas

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出自一致性框架,指出LM的缺陷源于单输出对独立评估的假设,可通过一致性优化解决,为开发通用一致性LLM提供思路。

中文摘要 AI 辅助

尽管语言模型(LM)的预训练与后训练流程日益复杂,仍存在诸多重要缺陷:模型过度依赖用户表述(“奉承性”)、逻辑泛化不完整,以及生成自信但错误的响应。本文认为,这些缺陷源于贯穿整个流程的建模假设:即行为可基于单输出对独立指定与评估。若不考量模型在不同输入下的响应关系,许多模型缺陷即便并非无法检测,也极难被发现。在这篇立场论文中,我们提出自一致性作为理解这些缺陷的框架。我们首先观察到,各类旨在改进LM行为特定方面的技术——涵盖对抗鲁棒性、事实连贯性等不同属性——均可被理解为通用“一致性优化”流程的特例,且可通过一套标准优化工具解决。随后,我们概述了通过优化一致性可实现的一系列新模型属性,并在最后探讨开发通用一致性LLM的意义,包括其将具备的能力及引发的质疑。

英文摘要

Despite ever-increasing sophistication in language model (LM) pre- and post-training pipelines, many important failures persist: models overcondition on user framing ("sycophancy"), exhibit incomplete logical generalization, and produce confident but incorrect responses. We argue that these failures arise from a modeling assumption permeating all aspects of the pipeline: that behavior can be specified and evaluated independently on single-output pairs. Many model failures are difficult, if not impossible, to detect without reasoning about relationships between a model's responses across inputs. In this position paper, we propose self-consistency as a framework for understanding these failures. We first observe that a wide variety of techniques designed to improve specific aspects of LM behavior-targeting properties as diverse as adversarial robustness and factual coherence-can be understood as special cases of a common "consistency optimization" procedure and addressed with a standard set of optimization tools. We next outline a set of new model properties that could be achieved by optimizing for consistency, and conclude with a discussion of what it would mean to develop generally consistent LMs, including the capabilities they would enable and the objections they raise.

补充信息

↑