AI 中文总结
本研究针对LLM表面对齐导致的情境脆弱问题,提出根基对齐框架,通过多维度评估缺陷并开发动态控制方法,实现锚定上下文与生成的智能体。
AI 中文摘要
当前大语言模型(LLMs)的对齐调优范式优先关注表面层面的行为——流畅性、安全性和语气一致性。尽管这类范式在日常聊天中效果尚可,但本研究认为,这种表面对齐掩盖了模型缺乏根基的问题,导致模型虽在风格上表现自信,却在情境中十分脆弱。我们提出了一种根基对齐(Grounded Alignment)框架,分析模型如何处理上下文(输入)和构建生成内容(输出),随后将这些根基化的行为与人类需求对齐。首先,我们评估情境根基方面的失败:SitTest测试显示,尽管模型拥有大上下文窗口,但最先进的模型仍难以对不断变化的环境保持一致的“心智模型”;ReCode进一步表明,模型依赖表面启发式而非深层句法依赖:它们“读取”大量历史信息,却未真正“理解”不断演变的情境。其次,我们评估生成根基方面的问题:我们引入分支因子(Branching Factor, BF)来映射LLM的生成过程,发现标准对齐调优将该生成空间压缩至过早的风格崩溃;事后分析还显示,模型常无法理解自身的生成内容。最后,我们提出用于根基交互的动态控制方法:AI Realtor展示了通过上下文工程弥补情境根基不足的方案;基础对齐模型协作(Base-Aligned Model Collaboration)将探索过程与风格约束解耦;我们还提出了用于可验证强化学习的退火采样(Annealed Sampling),并将这些思路应用于成瘾支持领域,其中模型生成的合理化内容为高风险领域提供了沟通接口。总体而言,本研究超越表面对齐,迈向同时锚定上下文与生成内容的智能体。
英文摘要
The current alignment tuning paradigm for Large Language Models (LLMs) prioritizes surface-level behaviors -- fluency, safety, and tonal consistency. While effective for casual chat, this thesis argues that such surface alignment masks a lack of grounding, creating models that are stylistically confident but situationally brittle. We propose a framework of Grounded Alignment, analyzing how models process context (Input) and structure generation (Output), then aligning these grounded behaviors to human needs. First, we evaluate failures in Situational Grounding. SitTest shows that despite large context windows, state-of-the-art models struggle to maintain a consistent "mental model" of a changing environment. ReCode further shows that models rely on surface heuristics rather than deep syntactic dependencies: they "read" extensive histories without truly "understanding" the evolving situation. Second, we evaluate Generative Grounding. We introduce the Branching Factor (BF) to map LLM generation, finding that standard alignment tuning constricts this landscape into premature stylistic collapse. Hindsight further shows that models often fail to understand their own generations. Finally, we propose Dynamic Control for grounded interaction. AI Realtor demonstrates context engineering to compensate for poor situational grounding. Base-Aligned Model Collaboration decouples exploration from stylistic constraints. We also present Annealed Sampling for verifiable reinforcement learning and apply these ideas to Addiction Support, where model-generated rationalization offers a communication interface for high-stakes domains. Collectively, this work moves beyond surface alignment toward agents anchored in both context and generation.
CommentsPhD Thesis submitted to UChicago (https://knowledge.uchicago.edu/records/5jrnc-nvp13). Update Branching Factor, AI Realtor, and BACo to the latest versions