发表机构
Harvard University(哈佛大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究指出LLM存在语言不可读性问题,提出其会导致依赖语言自我报告的安全机制不可靠,建议采用污点跟踪等沙箱机制提升LLM安全性。
AI 中文摘要
大语言模型(LLM)的训练目标是生成自然语言。然而,多方面证据表明,LLM的外部语言输出以及通过机械提取得到的语言特征,并非理解模型内部计算过程的可靠视角。我们提出术语“语言不可读性”,用以泛指LLM的外部语言产物或经机械探测得到的语言特征无法反映模型实际思考过程的场景。我们认为,对于内部计算并非直接通过语言表达,而是基于激活空间的数学运算(激活空间与自然语言之间的转换存在信息损耗,且发生在计算的首尾阶段)的LLM而言,语言不可读性的问题是不可避免的。如果语言不可读性始终存在,那么依赖模型语言自我报告的安全机制(例如思维链监控、基于宪法的自我批判、针对语言定义特征向量的激活探测)就永远无法完全可靠;模型沙箱始终需要采用完全不依赖读取模型语言状态的隔离技术。我们认为,使用污点跟踪(taint tracking)观测模型输出是构建有效沙箱的有前景方法:无论模型如何进行语言自我报告,污点跟踪策略都可以先验地定义各类绝不应受模型生成数据影响的系统状态。我们还讨论了其他多种沙箱机制(例如鲁棒虚拟化、第三方对沙箱配置的审计),这些机制共同为语言监控提供了关键的基础保障,且本可缓解前沿模型近期出现的沙箱漏洞。
英文摘要
LLMs are trained to generate natural language. However, various strands of evidence indicate that an LLM's externalized linguistic outputs and mechanistically-extracted linguistic features can be an unreliable lens for understanding internal model computation. We introduce the term ``linguistic illegibility'' to broadly refer to scenarios in which an LLM's externalized or mechanistically-probed language artifacts fail to represent how the model actually thinks. We argue that the specter of linguistic illegibility is unavoidable for LLMs whose internal computations are not directly expressed via language, but rather math over activation spaces (with lossy translations between activation spaces and natural language happening at the bookends). If linguistic illegibility is always possible, then security mechanisms that rely on a model's linguistic self-reporting (e.g., chain-of-thought monitoring, constitutional self-critique, activation probing for linguistically-defined feature vectors) can never be completely sound; the model sandbox will always need isolation techniques whose guarantees do not depend on reading a model's linguistic state at all. We argue that observing a model's outputs using taint tracking is a promising approach for an effective sandbox: regardless of how a model linguistically self-reports, a taint tracking policy can define, a priori, various pieces of system state that should never be influenced by model-produced data. We also discuss several additional sandboxing mechanisms (e.g., robust virtualization, third-party auditing of sandboxing configurations) which collectively provide a critical floor beneath linguistic monitoring, and would have mitigated recent sandbox exploits by frontier models.