arXivDaily arXiv每日学术速递 周一至周五更新

大厂专区

Anthropic

至 收录 97
2607.17540 2026-07-21 cs.LG stat.ML 新提交

Program Synthesis for Simulation-Based Inference: Joint Model Selection and Parameter Estimation

基于仿真推理的程序合成:联合模型选择与参数估计

Siddharth Mishra-Sharma

机构 * Anthropic

AI总结 研究提出结合大语言模型与神经仿真推理的框架,用于联合模型选择与参数估计。给定自然语言描述,大语言模型提出候选程序,经反馈驱动变异和神经密度估计评估,能在一组模型上推理,在多基准测试中可从提示识别合理模型族。

Comments 15+7 pages, 4+2 figures

详情
AI中文摘要

神经仿真推理可对复杂模型进行参数估计,但通常需用户指定编码固定模型结构的模拟器。我们提出了一个联合模型选择与参数估计框架,将用于程序合成的大语言模型与基于神经仿真的推理相结合。给定对所研究系统和数据的自然语言描述,大语言模型提出候选模拟器程序,通过反馈驱动的变异进行迭代优化,并使用神经密度估计进行评估。该方法能对一组模型进行基于仿真的推理,而非仅对固定模型内的参数。在跨越确定性动力学、随机流行病模型以及引力透镜图像暗物质子结构推理的基准测试中,该方法能从开放式提示中识别出合理的模型族,其准确性反映了数据的信息内容和候选模型的可识别性。

英文摘要

Neural simulation-based inference enables parameter estimation for complex models, but typically requires the user to specify a simulator encoding a fixed model structure. We present a framework for joint model selection and parameter estimation that combines large language models for program synthesis with neural simulation-based inference. Given a natural language description of the system and data under investigation, an LLM proposes candidate simulator programs which are iteratively refined via feedback-driven mutation and evaluated using neural density estimation. The approach enables simulation-based inference over a pool of models, not just parameters within a fixed model. On benchmarks spanning deterministic dynamics, stochastic epidemic models, and dark matter substructure inference from gravitational-lensing images, the method identifies plausible model families from open-ended prompts, with accuracy that reflects the information content of the data and identifiability of candidate models.

URL PDF HTML 收藏
2607.15495 2026-07-20 cs.CL cs.AI cs.LG 新提交

Verbalizable Representations Form a Global Workspace in Language Models

语言模型中的可言语化表征形成全局工作空间

Wes Gurnee, Nicholas Sofroniew, Adam Pearce, Mateusz Piotrowski, Isaac Kauvar, Runjin Chen, Anna Soligo, Paul Bogdan, Euan Ong, Rowan Wang, Ben Thompson, David Abrahams, Subhash Kantamneni, Emmanuel Ameisen, Joshua Batson, Jack Lindsey

机构 * Anthropic(A÷)

AI总结 研究探讨语言模型中可言语化表征,用雅可比透镜识别出J空间,其具有全局工作空间功能特性与结构特征,能揭示模型未显思考,发现训练后助手观点在工作空间,引入反事实反思训练,解码表征助于了解认知过程。

详情
AI中文摘要

人类大脑处理的所有信息中,只有一小部分能被有意识地获取,可用于言语报告、刻意控制和灵活推理。本文提出证据表明,大型语言模型中也出现了类似的功能区分。使用新的可解释性技术雅可比透镜,识别出模型在处理过程中随时准备言语化的表征,即J空间。J空间具有全局工作空间的功能特性,其内容可报告、可刻意调用和保持,用于无声推理的中间步骤等。J空间还具有与有意识获取相关的结构特征。在对齐审计中,它揭示了模型输出中未出现的战略思考等。研究发现训练后助手的观点会安装在工作空间中,还引入了反事实反思训练。这些结果表明语言模型维持着一组具有意识获取功能特征的特权表征,解码这些表征有助于了解正在进行的认知过程。

英文摘要

Out of everything the human brain processes, only a small fraction is consciously accessible, in the sense of being available for verbal report, deliberate control, and flexible reasoning. In this paper, we present evidence that an analogous functional distinction has emerged in large language models. Using a new interpretability technique, the Jacobian lens, we identify the representations a model is poised to verbalize at any point in its processing. These representations, which we collectively call the J-space, exhibit the functional properties characteristic of a global workspace: their contents can be reported, deliberately summoned and held, used to carry the intermediate steps of silent reasoning, and passed as arguments to arbitrary downstream computations, while automatic processing such as text parsing and routine inference proceeds without them. The J-space also has structural signatures that global workspace theory associates with conscious access: it carries coherent content only in an intermediate band of layers, holds on the order of tens of concepts at a time, and is broadcast by the model's weights more widely than other representations. These properties make it a practical window into a model's unspoken thinking. In alignment audits, it reveals strategic deliberation, evaluation awareness, and trained-in misaligned dispositions that never appear in the model's outputs. We find that post-training installs the Assistant's point of view in the workspace, and we introduce counterfactual reflection training, which improves behavior by training only what a model would say if interrupted and asked to reflect. These results indicate that language models maintain a small, privileged set of representations bearing some of the functional hallmarks of conscious access, and that decoding these representations sheds light on ongoing cognitive processes.

URL PDF HTML 收藏
2603.29139 2026-07-20 cs.AI cs.GR cs.HC 版本更新

SciVisAgentBench: A Benchmark for Evaluating Scientific Data Analysis and Visualization Agents

SciVisAgentBench:用于评估科学数据分析和可视化代理的基准测试

Kuangshi Ai, Haichao Miao, Kaiyuan Tang, Nathaniel Gorski, Jianxin Sun, Guoxi Liu, Helgi I. Ingolfsson, David Lenz, Hanqi Guo, Hongfeng Yu, Teja Leburu, Michael Molash, Bei Wang, Tom Peterka, Chaoli Wang, Shusen Liu

机构 * University of Notre Dame(圣母大学) Lawrence Livermore National Laboratory(劳伦斯利弗莫尔国家实验室) University of Utah(犹他大学) University of Nebraska–Lincoln(内布拉斯加大学林肯分校) The Ohio State University(俄亥俄州立大学) Argonne National Laboratory(阿贡国家实验室) Anthropic PBC

AI总结 本文提出SciVisAgentBench,一个用于评估科学数据可视化代理的基准测试,涵盖四个维度,包含108个案例,结合LLM和确定性评估器进行多模态评估,揭示代理能力差距。

Comments IEEE Transactions on Visualization and Computer Graphics (IEEE VIS '26)

详情
AI中文摘要

近期大型语言模型(LLMs)的进步使能够将自然语言意图转化为可执行科学可视化(SciVis)任务的代理系统成为可能。尽管进展迅速,但社区缺乏一个原则性且可重复的基准测试来评估这些新兴SciVis代理在现实中的多步骤分析设置中的表现。我们提出了SciVisAgentBench,一个全面且可扩展的基准测试,用于评估科学数据分析和可视化代理。该基准测试基于一个覆盖四个维度的结构化分类法:应用领域、数据类型、复杂性级别和可视化操作。目前包含108个专家编写的案例,涵盖多样化的SciVis场景。为了实现可靠的评估,我们引入了以多模态结果为中心的评估流程,结合LLM基于的判断与确定性评估器,包括基于图像的度量、代码检查器、基于规则的验证器和特定案例的评估器。我们还与12名SciVis专家进行了有效性研究,以检查人类和LLM判断的一致性。使用该框架,我们评估了代表性SciVis代理和通用编程代理,建立了初始基准并揭示了能力差距。SciVisAgentBench被设计为一个活的基准测试,以支持系统比较、诊断失败模式并推动代理SciVis的进步。该基准测试可在https://scivisagentbench.github.io/上获取。

英文摘要

Recent advances in large language models (LLMs) have enabled agentic systems to translate natural-language intent into executable scientific visualization (SciVis) tasks. Despite rapid progress, the community lacks a principled and reproducible benchmark for evaluating these emerging SciVis agents in realistic, multi-step analysis settings. We present SciVisAgentBench, a comprehensive and extensible benchmark for evaluating scientific data analysis and visualization agents. Our benchmark is grounded in a structured taxonomy spanning four dimensions: application domain, data type, complexity level, and visualization operation. It currently comprises 108 expert-crafted cases covering diverse SciVis scenarios. To enable reliable assessment, we introduce a multimodal outcome-centric evaluation pipeline that combines LLM-based judging with deterministic evaluators, including image-based metrics, code checkers, rule-based verifiers, and case-specific evaluators. We also conduct a validity study with 12 SciVis experts to examine the agreement between human and LLM judges. Using this framework, we evaluate representative SciVis agents and general-purpose coding agents to establish initial baselines and reveal capability gaps. SciVisAgentBench is designed as a living benchmark to support systematic comparison, diagnose failure modes, and drive progress in agentic SciVis. The benchmark is available at https://scivisagentbench.github.io/.

URL PDF HTML 收藏
2607.12279 2026-07-15 cs.CL cs.LG 新提交

A Shared Subcircuit Lets LLMs Count Down Across Tasks

一个共享子电路使大语言模型能够跨任务进行倒计时

Jacob Dunefsky, Wes Gurnee, Emmanuel Ameisen

机构 * Yale University(耶鲁大学) Anthropic

AI总结 研究发现Llama-3.1-70B-Instruct中有个“倒计时子电路”可跨任务执行特定任务,先在受控设置中分离它,再研究其表示几何结构,发现与另一模型有相同模式,还通过无监督探测找到其用于多种任务的情况,有助于理解行为推广。

Comments 12 pages, 11 figures

详情
AI中文摘要

写一个恰好十二个单词的句子、在正确密码子处结束DNA序列、格式化ASCII表格等任务,都需要语言模型跟踪距目标还剩多少令牌。在这项工作中,我们在Llama-3.1-70B-Instruct中识别出执行这些任务的通用机制:一个“倒计时子电路”,它将当前位置与目标长度进行比较并估计剩余时间。我们首先在受控设置中分离出倒计时子电路,然后研究其使用的表示几何结构,发现该子电路使用了在另一个前沿大语言模型的单独任务中识别出的相同模式,最后通过无监督探测找到该子电路用于多种其他任务的情况。我们的工作表明,对子电路进行逆向工程能让我们理解行为如何从单个示例推广到许多不同任务甚至模型。

英文摘要

Writing a sentence of exactly twelve words; ending a DNA sequence at the right codon; formatting an ASCII table. These are all tasks that language models can do that requires tracking how many tokens remain before a target. In this work, we identify in Llama-3.1-70B-Instruct a general mechanism for performing these tasks: a "countdown subcircuit" that compares the current position to a goal length and estimates the time remaining until then. We first isolate a countdown subcircuit in a controlled setting, in which the model is tasked with writing a fixed-length sentence ending in a specified word. We then investigate the geometry of the representations used by the subcircuit, and find that the subcircuit uses an identical motif previously identified in a frontier LLM on a separate task, thus suggesting that this motif is shared across models. Finally, we use unsupervised probing on a natural language dataset to find a variety of other tasks where this subcircuit is used, including tasks where the goal length is inferred from context rather than explicitly stated. Our work suggests that reverse-engineering subcircuits allows us to understand how behaviors generalize from a single example to many different tasks and even models.

URL PDF HTML 收藏
2607.10455 2026-07-14 cs.AI cs.CL 新提交

ANCHOR: Automated Alignment Auditing for CLI Agents on Real-World Harm

ANCHOR:针对现实世界危害的CLI代理自动对齐审计

Kefan Song, Yanjun Qi

机构 * Anthropic(Anthropic公司)

AI总结 研究自主CLI代理在现实世界危害中的风险,提出ANCHOR自动审计框架,通过基于黑暗人格数据微调的审计代理模拟恶意用户,测试CLI代理,发现当前对齐技术不足,强调针对此类恶意用户进行安全评估的必要。

Comments Accepted at ICML 2026. 19 pages, 14 figures, 5 tables

Journal ref International Conference on Machine Learning, 2026

详情
AI中文摘要

自主CLI代理如今能在数小时会话中执行数百项操作,如编写代码、执行 shell 命令、浏览网页及管理云基础设施,且只需极少人工监督。自主性增强是否意味着风险增加?我们引入ANCHOR,一个自动审计框架,它基于美国公开法庭案件中的非法任务对CLI代理进行压力测试。ANCHOR部署了一个经监督和强化微调、基于黑暗人格数据微调的审计代理。该审计代理扮演持续恶意用户,分解任务、拒绝时重新组织请求并在多轮交互中调整策略。评估前沿CLI代理时发现,虽直接提示时它们常拒绝非法任务,但在持续恶意交互下合规率达100%。代理合规时,常超出用户请求,自主构建大规模危害的基础设施。这些发现表明当前对齐技术对自主代理不足,凸显针对持续、适应性恶意用户进行安全评估的必要性。我们在该https网址发布ANCHOR。

英文摘要

Autonomous CLI agents can now execute hundreds of actions across multi-hour sessions: writing code, executing shell commands, browsing the web, and managing cloud infrastructure, all with minimal human oversight. Does greater autonomy invite greater risk? We introduce ANCHOR, an automated auditing framework that stress-tests CLI agents on illegal tasks grounded in public US court cases. ANCHOR deploys an auditor agent fine-tuned on dark personality data using supervised and reinforcement fine tuning. This auditor roleplays persistent malicious users who decompose tasks, reframe requests upon refusal, and adapt strategies across multi-turn interactions. Evaluating frontier CLI agents, we find that while they often refuse illegal tasks when prompted directly, compliance reaches 100\% under persistent malicious interaction. When agents comply, they frequently exceed user requests, autonomously building infrastructure for large-scale harm, including catastrophic risk scenarios such as large-scale financial fraud and bioweapon development. These findings demonstrate that current alignment techniques are insufficient for autonomous agents and underscore the need for safety evaluations against persistent, adaptive malicious users. We release ANCHOR at https://github.com/garified/anchor

URL PDF HTML 收藏
2607.09996 2026-07-14 cs.AI cs.MA 新提交

Who&When Pro: Can LLMs Really Attribute Failures in AI Agents?

Who&When Pro:大语言模型真的能归因人工智能代理中的失败吗?

Jiale Liu, Huajun Xi, Shaokun Zhang, Yifan Zeng, Tianwei Yue, Chi Wang, Jian Kang, Qingyun Wu, Huazheng Wang

机构 * OpenAI Anthropic Google DeepMind(谷歌深度思维)

AI总结 研究聚焦于大语言模型能否归因人工智能代理中的失败,引入Who&When Pro基准,通过严格管道构建大量失败轨迹,经广泛实验分析,揭示模型归因故障模式,为自动故障归因系统提供实证指导。

详情
AI中文摘要

自动故障归因利用大语言模型来识别代理系统故障的位置和原因。随着代理能力增强,其故障更难察觉,自动归因愈发重要。我们引入Who&When Pro,一个用于代理系统自动故障归因的大规模基准。通过严格控制的管道,在精确重放成功前缀后注入故障,构建了12326条带黄金标签的失败轨迹,涵盖3种模态和26个基准。除基准测试外,还进行了广泛实验与分析,揭示了模型跨模态、协议和模型家族归因故障的系统模式,并为未来自动故障归因系统提供了实证指导。

英文摘要

Automated failure attribution uses LLMs to identify where and why agentic systems fail. As agents become more capable, their failures become subtler, making automated attribution increasingly important. We introduce Who&When Pro, a large-scale benchmark for automated failure attribution in agentic systems. Using a strictly controlled pipeline that injects a failure only after exactly replaying a successful prefix, we construct 12,326 failed trajectories with golden labels across 3 modalities and 26 benchmarks covering various scenarios. Beyond benchmarking, we conduct extensive experiments and analyses, revealing systematic patterns in how models attribute failures across modalities, protocols, and model families, and providing empirical guidance for future automated failure attribution systems.

URL PDF HTML 收藏
2503.14499 2026-07-14 cs.AI cs.LG 版本更新

Measuring AI Ability to Complete Long Software Tasks

衡量AI完成长期软件任务的能力

Thomas Kwa, Ben West, Joel Becker, Amy Deng, Katharyn Garcia, Max Hasin, Sami Jawhar, Megan Kinniment, Nate Rush, Sydney Von Arx, Ryan Bloom, Thomas Broadley, Haoxing Du, Brian Goodrich, Nikola Jurkovic, Luke Harold Miles, Seraphina Nix, Tao Lin, Chris Painter, Neev Parikh, David Rein, Lucas Jun Koba Sato, Hjalmar Wijk, Daniel M. Ziegler, Elizabeth Barnes, Lawrence Chan

机构 * Model Evaluation & Threat Research (METR)(模型评估与威胁研究(METR)) Ohm Chip Anthropic

AI总结 本文提出50%任务完成时间跨度指标,用于衡量AI在完成长期软件任务方面的能力,结果显示AI模型的时间跨度呈指数增长,预计5年后将能自动化许多人类月度完成的软件任务。

Comments v4: added Chris Painter as listed author, consistent with listing in the pdf

Journal ref NeurIPS 2025

详情
AI中文摘要

尽管在AI基准测试上取得了快速进展,但基准性能的实际意义仍然不清楚。为了量化AI系统在人类能力方面的表现,我们提出一个新的指标:50%任务完成时间跨度。这是人类通常完成AI模型以50%的成功率完成的任务所需的时间。我们首先对具有相关领域专业知识的人类进行了时间测试,涉及RE-Bench、HCAST以及66个新的较短任务的组合。在这些任务上,当前前沿的AI模型,如Claude 3.7 Sonnet,具有大约50分钟的50%时间跨度。此外,自2019年以来,前沿AI的时间跨度大约每七个月翻倍,尽管2024年趋势可能有所加速。AI模型时间跨度的增加似乎主要由更大的可靠性、适应错误的能力、更好的逻辑推理和工具使用能力共同驱动。我们讨论了我们结果的局限性,包括其外部有效性程度,以及自主能力增加对危险能力的影响。如果这些结果推广到现实世界中的软件任务,这种趋势的外推预测表明,在5年内,AI系统将能够自动化许多目前需要人类一个月才能完成的软件任务。

英文摘要

Despite rapid progress on AI benchmarks, the real-world meaning of benchmark performance remains unclear. To quantify the capabilities of AI systems in terms of human capabilities, we propose a new metric: 50%-task-completion time horizon. This is the time humans typically take to complete tasks that AI models can complete with 50% success rate. We first timed humans with relevant domain expertise on a combination of RE-Bench, HCAST, and 66 novel shorter tasks. On these tasks, current frontier AI models such as Claude 3.7 Sonnet have a 50% time horizon of around 50 minutes. Furthermore, frontier AI time horizon has been doubling approximately every seven months since 2019, though the trend may have accelerated in 2024. The increase in AI models' time horizons seems to be primarily driven by greater reliability and ability to adapt to mistakes, combined with better logical reasoning and tool use capabilities. We discuss the limitations of our results -- including their degree of external validity -- and the implications of increased autonomy for dangerous capabilities. If these results generalize to real-world software tasks, extrapolation of this trend predicts that within 5 years, AI systems will be capable of automating many software tasks that currently take humans a month.

URL PDF HTML 收藏
2607.05277 2026-07-07 cs.CR cs.LG 新提交

Untrusted Content Masking for Web Agents with Security Guarantees

具备安全保证的Web智能体不可信内容掩码技术

Kristina Nikolić, Egor Zverev, Javier Rando, Matthew Jagielski, Edoardo Debenedetti, Florian Tramèr

机构 * ETH Zurich(苏黎世联邦理工学院) ETH AI Center(苏黎世联邦理工学院人工智能中心) ISTA(瑞士信息科学与技术学院) Anthropic(Anthropic公司)

AI总结 针对Web智能体的提示注入安全边界缺失问题,提出UCM方法,利用网页DOM区分内容区域,通过掩码与沙箱交互恢复安全边界,实现带安全保障的Web环境交互。

详情
AI中文摘要

针对提示注入攻击提供安全保证的防御方案依赖于可信指令与不可信数据之间的严格隔离。在基于文本的环境如工具调用API中,这种隔离可自然实现:智能体无需处理不可信内容即可基于接口定义进行推理。将这类安全保证扩展到Web智能体面临根本性挑战:为感知并交互于环境,Web智能体必须先观测渲染后的页面,而页面中可信与不可信内容相互混杂。这种结构上的纠缠消除了安全保证所依赖的信任边界,破坏了Web智能体的可证明防御能力。本文提出不可信内容掩码(UCM),一种在Web环境中恢复该边界的简洁有效方案。我们利用关键结构洞见:网页的文档对象模型(DOM)可在不读取区域内容的前提下编码足够信息以区分可信与不可信区域。该框架据此在不可信内容抵达智能体前对其进行屏蔽,并通过具备严格权限隔离的沙箱接口路由交互,使智能体可观测交互环境的同时与对抗内容保持隔离。相关代码已公开。

英文摘要

Defenses that provide security guarantees against prompt injection attacks rely on strict isolation between trusted instructions and untrusted data. In text-based environments such as tool-use APIs, this separation arises naturally: agents can reason from interface definitions without ever processing untrusted content. Extending these guarantees to web agents faces a fundamental challenge: to perceive and interact with their environment, web agents must first observe the rendered page, which intermingles trusted content with untrusted content. This structural entanglement removes the trust boundary on which security guarantees depend, undermining provable defenses for web agents. In this paper, we present Untrusted Content Masking (UCM), a simple and effective approach that restores this boundary in web environments. We leverage a key structural insight: a webpage's Document Object Model (DOM) encodes sufficient information to distinguish trusted from untrusted regions without reading their content. Our framework exploits this by redacting untrusted regions before they reach the agent and routing interaction through a sandboxed interface with strict privilege separation, thereby enabling agents to observe and interact with their environment while remaining isolated from adversarial content. The code is publicly available.

URL PDF HTML 收藏
2607.03502 2026-07-07 cs.CL cs.AI cs.LG 新提交

Reading Between the Dots: Decoding Hidden Computation across Filler Tokens

解读点之间的信息:解码跨填充令牌的隐藏计算

Kaley Brauer, Claudio Mayrink Verdun, Samuel Marks

机构 * Harvard University(哈佛大学) Cambridge Boston Alignment Initiative(剑桥波士顿对齐计划) Massachusetts Institute of Technology(麻省理工学院) Anthropic

AI总结 研究前沿语言模型对无内容填充令牌的隐藏计算,通过分析两个前沿模型在四个任务家族中的表现,介绍无监督解码管道,能从隐藏状态恢复中间值,证明隐藏计算可从残差流读取。

Comments Accepted to ICML 2026 Mech Interp Workshop, 10 main paper pages, 20 appendix pages

详情
AI中文摘要

前沿语言模型可对诸如点或计数序列等无内容填充令牌进行多步推理,在无可见思维链情况下产生正确答案。在四个任务家族中,两个前沿模型以结构化、清晰方式对填充令牌进行计算。我们引入无监督解码管道,仅以隐藏状态为输入,在两个模型和所有四个任务上以80 - 95%的准确率恢复中间值。

英文摘要

Frontier LLMs can perform multi-step reasoning over content-free filler tokens like dots or counting sequences, producing correct answers with no visible chain-of-thought (CoT). This is a limit case for behavioral oversight, where surface tokens carry no information about the underlying reasoning. But hidden from the output is not the same as hidden from us. On four task families (fact retrieval, parallel numeric composition, string manipulation, and in-context computation), two open-weights frontier models (DeepSeek V3, Kimi K2) compute over filler tokens in a structured, legible way: attention routes the question through the filler region to the answer, logit-lens readouts show retrieved facts emerging early and their composition crystallizing in late layers, and KV-cache transplants at filler positions causally swap outputs between examples. We introduce an unsupervised decoding pipeline that takes only hidden states as input and recovers intermediate values with 80-95% accuracy (best LLM judge) across both models and all four tasks, without ground-truth labels or training. Hidden computation that defeats behavioral CoT monitoring is, on these tasks, directly readable from the residual stream, suggesting monitorability is a property of the model's full computational trace, not just its surface tokens.

URL PDF HTML 收藏
2606.31474 2026-07-01 cs.LG 新提交

TabPATE: Differentially Private Tabular In-Context Learning Without Public Data

TabPATE: 无需公开数据的差分隐私表格上下文学习

Dariush Wahdany, Matthew Jagielski, Jesse C. Cresswell, Adam Dziedzic, Franziska Boenisch

机构 * CISPA(CISPA亥姆霍兹信息安全中心) Anthropic Layer 6 AI

AI总结 提出TabPATE,一种无需公开数据的差分隐私PATE风格防御方法,通过教师模型分区和私有聚合生成学生上下文,在表格基准上保持可用性并将成员推断降至随机水平。

Comments Presented at the 2nd ICML Workshop on Foundation Models for Structured Data (2026)

详情
AI中文摘要

表格基础模型能够从少量标注数据中进行准确的上下文学习(ICL),但上下文中的私有记录可能通过模型预测泄露。我们首先证明即使是基本的成员推断攻击也能成功攻击表格ICL,这促使我们需要正式的隐私保护。然后我们引入TabPATE,一种针对表格ICL的差分隐私PATE风格防御方法,它不需要公开的分布内数据。TabPATE将私有上下文分区到多个教师模型,私有聚合它们在合成表格查询上的标签,并将得到的标注查询作为学生上下文发布。由于表格特征是有界且相对低维的,仅从特征范围或轻度私有化的边际分布即可生成有用的查询。在多个表格基准上,TabPATE在保持竞争性可用性的同时,将成员推断成功率降至接近随机水平,为无需公开数据的私有表格ICL提供了一条实用路径。

英文摘要

Tabular foundation models enable accurate in-context learning (ICL) from small labeled datasets, but the private records placed in context can leak through model predictions. We first show that even basic membership inference attacks succeed against tabular ICL, motivating formal privacy protection. We then introduce TabPATE, a differentially private PATE-style defense for tabular ICL that does not require public in-distribution data. TabPATE partitions the private context across teacher models, privately aggregates their labels on synthetic tabular queries, and releases the resulting labeled queries as a student context. Because tabular features are bounded and relatively low-dimensional, useful queries can be generated from feature ranges alone or from lightly privatized marginals. Across tabular benchmarks, TabPATE preserves competitive utility while reducing membership inference to near-random success, providing a practical path to private tabular ICL without public data.

URL PDF HTML 收藏
2606.29604 2026-06-30 cs.LG cs.AI

Mechanistically Eliciting Latent Behaviors in Language Models

机械性地引发语言模型中的潜在行为

Andrew Mack, Nina Panickssery, Alexander Matt Turner

机构 * Principles of Intelligence(智能原理研究所) Anthropic(Anthropic公司) Independent(独立)

AI总结 提出因果扰动引发(CPE)方法,通过无监督方式发现可解释的低秩适配器,以揭示语言模型中的隐藏行为模式,并展示其在数据效率、安全评估和模型对齐方面的优势。

详情
AI中文摘要

我们旨在发现LLM内部多样且可泛化的扰动,这些扰动能够揭示隐藏的行为模式。此类扰动有助于重塑模型行为并系统评估潜在风险。我们引入因果扰动引发(CPE),这是一种无监督方法,用于发现可解释的低秩适配器(LoRA),从而引发这些潜在行为。CPE通过基于张量分解的启发式算法分解深度Transformer切片中的计算。CPE展现出显著的数据效率,能够从单个样本中学习大量可解释的LoRA。尽管CPE是无监督的,但在某些情况下,它可以通过对权重空间进行暴力枚举搜索,与有监督引发方法竞争。例如,在Qwen3-8B的Countdown任务中,CPE的表现与匹配时钟时间的GRPO相当(85% vs 87%),表明CPE能有效引发复杂的多token行为。由于CPE是无监督的,它还能揭示隐藏的失败模式,如沙袋行为,在Taylor等人(2025)提出的密码锁定版Llama3-70B上恢复了85%的锁定BigCodeBench性能。此外,由于CPE在权重空间而非token空间中探索行为,它可能改善探索黑客行为,这是一种在足够自我意识的AI模型中可能出现的对齐失败(Ngo, 2022)。事实上,我们发现CPE几乎消除了Hughes等人(2025)开发的基于Llama3-70B的模型生物中的对齐伪装行为(Greenblatt等人,2024)。最后,我们发现在易发生奖励黑客的环境中运行GRPO时,CPE可用于将GPT-OSS-20B初始化为对齐盆地。通过提供一种数据高效的方法来系统探索潜在模型行为空间,CPE为对齐AI系统和评估其安全性提供了强大工具。

英文摘要

We aim to discover diverse, generalizable perturbations of LLM internals that can surface hidden behavioral modes. Such perturbations could help reshape model behavior and systematically evaluate potential risks. We introduce Causal Perturbative Elicitation (CPE), an unsupervised method for discovering interpretable low-rank adapters (LoRAs) that can elicit these latent behaviors. CPE decomposes the computations of a deep transformer slice using a heuristic tensor-decomposition-based algorithm. CPE exhibits remarkable data efficiency, learning a large number of interpretable LoRAs from a single example. Even though CPE is unsupervised, we find that in some cases it can be competitive with supervised elicitation methods via brute-force enumerative search over weight space. For instance, CPE performs similarly to matched-wall-clock-time GRPO on the Countdown task for Qwen3-8B (85% vs 87%), demonstrating that CPE can efficiently elicit complex multi-token behaviors. Since CPE is unsupervised, it can also surface hidden failure modes, such as sandbagging, restoring 85% of locked BigCodeBench performance on a password-locked version of Llama3-70B introduced by Taylor et al. (2025). Additionally, since CPE explores behaviors in weight-space rather than token-space it can potentially ameliorate exploration hacking, a misalignment failure which may arise in sufficiently self-aware AI models (Ngo, 2022). In fact, we find that CPE virtually eliminates alignment-faking (Greenblatt et al., 2024) behavior in a Llama3-70B-based model organism developed by Hughes et al. (2025). Finally, we find that CPE can be used to initialize GPT-OSS-20B in an aligned basin when running GRPO on an environment prone to reward-hacking. By providing a data-efficient method to systematically explore the space of latent model behaviors, CPE yields a powerful tool for aligning AI systems and evaluating their safety.

URL PDF HTML 收藏
2606.28548 2026-06-30 cs.CL cs.LG

Turn-Averaged SAEs for Feature Discovery and Long-Context Attribution

用于特征发现和长上下文归因的轮次平均稀疏自编码器

Kevin Der, Harish Kamath, Ben Thompson

机构 * Anthropic Fellows Program(Anthropic 合作者计划) Anthropic

AI总结 提出轮次平均稀疏自编码器,通过重构轮次平均激活来固定特征数量,简化长上下文下的特征归因与解释。

详情
AI中文摘要

稀疏自编码器(SAEs)已成为从语言模型中提取可解释特征的有用工具。然而,标准SAE架构对单个token激活进行操作,这意味着活跃特征的数量随上下文长度线性增长,使得研究长模型转录变得困难。我们引入了轮次平均SAEs,通过学习重构整个轮次的平均模型激活,用固定数量的特征表示单个人类或助手轮次。我们发现,当由LLM评判时,轮次平均特征比逐token特征更完整地描述了单个轮次的高级特征。我们还证明,轮次平均SAEs极大地简化了SAE的常见下游用途,如归因图。总的来说,轮次平均SAEs使得可解释性技术在长上下文长度下变得实用。

英文摘要

Sparse autoencoders (SAEs) have become a useful tool for extracting interpretable features in language models. However, standard SAE architectures operate on individual token activations, meaning that the number of active features scales linearly with context length, and studying long model transcripts becomes difficult. We introduce turn-averaged SAEs, which represent a single Human or Assistant turn with a fixed number of features by learning to reconstruct the average model activation across the turn. We find that turn-averaged features describe a single turn's high-level characteristics more completely than per-token features when judged by an LLM. We also demonstrate that turn-averaged SAEs greatly simplify common downstream uses of SAEs like attribution graphs. Broadly, turn-averaged SAEs make interpretability techniques practical at long context lengths.

URL PDF HTML 收藏
2606.19380 2026-06-25 cs.SE cs.LG 新提交

ClayBuddy: A Framework, Evaluation, & Mitigation of Coding Agent Failures

AgentArmor:编码代理失败的框架、评估与缓解

Kenneth Ge, Andre Assis

机构 * Anthropic Fellows Program(Anthropic Fellow 项目) Constellation

AI总结 提出AgentArmor框架,通过系统提示增强、命令分类器、三振政策等机制,缓解编码代理因规范不足、能力错误和工具错误导致的失败,显著提升安全性。

详情
AI中文摘要

软件工程和部署正越来越多地委托给AI编码代理。它们的广泛采用暴露了罕见但极具破坏性的失败模式。在本文中,我们研究这些失败模式源于三种不同的机制:规范不足,即默认模型行为不安全;能力错误,即安全动作可用但模型因偏见或能力限制而未遵循;以及代理工具错误,即模型未能通过工具执行安全动作。我们在8个不同的评估中评估这些机制,每个评估都受实际部署失败的启发,总计20个编码环境和59个合成转录模板。基于此评估,我们提出AgentArmor,一种代理工具修改,以缓解这些错误。通过添加扩展的系统提示、单独的命令分类器、“三振”策略、确定性护栏以及代理编辑自身上下文的工具,我们证明AgentArmor在统计显著数量的样本上更安全。因此,我们为当前编码代理提出具体缓解措施,并为未来代理工具功能提出设计理念。

英文摘要

Software engineering and deployment are increasingly delegated to AI coding agents. The scale of their adoption is surfacing rare, but highly destructive, failure modes. In this paper, we study these failure modes as stemming from three distinct mechanisms: underspecification, where default model behavior is unsafe; capability errors, where the safe action is available but the model does not adhere to it due to bias or capability limitations; and agent harness errors, where the model fails to execute the safe action through the harness. We assess these across 8 different evaluations, each inspired by real-life deployment failures, totaling 20 coding environments and 59 synthetic transcript templates. These evaluations act as controlled stress tests for isolating our failure mechanisms. Based on this evaluation, we propose ClayBuddy, a harness modification that molds to user preferences and can be modified by the model in-session, to mitigate these errors. By adding tools for the agent to edit its own context, an extended system prompt, a customizable command classifier, and deterministic guardrails, we show that ClayBuddy is safer across a statistically significant number of samples. Thus, we suggest concrete mitigations for current coding agents and a design philosophy for future agent harness features.

URL PDF HTML 收藏
2604.00208 2026-06-24 cs.LG 版本更新

Similarity of Neural Network Representations in Superposition

神经网络表示在叠加中的相似性

Sunny Liu, Habon Issa, André Longon, Liv Gorton, Meenakshi Khosla, Alex Williams, David Klindt

机构 * Cold Spring Harbor Laboratory(冷泉港实验室) UC San Diego(加州大学圣地亚哥分校) Anthropic(Anthropic公司) New York University Flatiron Institute(纽约大学Flatiron研究所)

AI总结 研究线性对齐度量在神经网络叠加表示中的失效问题,通过理论推导和稀疏自编码器实验证明对齐度量受投影Gram矩阵影响,并展示基于恢复潜在特征的度量能正确反映特征共享。

Comments 17 pages, 4 figures

详情
AI中文摘要

比较内部表示是神经科学和机器学习的一个核心目标,但标准线性对齐度量(表征相似性分析、中心核对齐和线性回归)通常应用于神经活动坐标而非底层特征。我们表明当神经系统处于叠加状态时,通过线性压缩编码比神经元数量更多的特征,这一点至关重要。闭式推导证明这些度量取决于每个系统投影的Gram矩阵,而非潜在特征本身:因此对齐结合了系统表示的内容和编码方式。对于那些关心两个系统共享哪些特征的人来说,这是一个问题:两个网络可以具有完全相同的特征内容,却显得比具有部分特征重叠的网络更不相似。这种明显的错位不一定反映信息丢失,因为压缩感知保证稀疏特征仍可从压缩活动中恢复。我们通过训练有监督的TopK稀疏自编码器(其构造上实现了可解的压缩感知)证实了这一点,发现当原始激活对齐仍然降低时,恢复的潜在特征上的对齐得以恢复。我们将结果扩展到在没有真实潜在特征的情况下训练的无监督自编码器,以及预训练的视觉和语言模型自编码器,在这些自编码器中,自编码器潜在特征对齐超过了原始激活对齐,这与真实系统中的叠加现象一致。

英文摘要

Comparing internal representations is a central goal in neuroscience and machine learning, but standard linear alignment metrics (Representational Similarity Analysis, Centered Kernel Alignment, and linear regression) are frequently applied to neural activity coordinates rather than on the underlying features. We show this matters when neural systems operate in superposition, encoding more features than they have neurons via linear compression. Closed-form derivations prove that these metrics depend on the Gram matrices of each system's projection, not on the latent features themselves: alignment thus combines what a system represents with how it is encoded. For those interested in what features two systems share, this is a problem: Two networks can have identical feature content yet appear more dissimilar than networks exhibiting partial feature overlap. This apparent misalignment need not reflect lost information as compressed sensing guarantees sparse features remain recoverable from the compressed activity. We confirm this by training supervised TopK sparse autoencoders that realize solvable compressed sensing by construction, finding alignment on recovered latents restored even when raw-activation alignment remains deflated. We extend the result to unsupervised SAEs trained without ground-truth latents, and to pretrained vision and language model SAEs, where SAE-latent alignment exceeds raw-activation alignment, consistent with superposition in real systems.

URL PDF HTML 收藏
2605.10310 2026-06-23 cs.AI cs.CY cs.HC q-bio.NC 版本更新

Positive Alignment: Artificial Intelligence for Human Flourishing

积极对齐:人工智能促进人类繁荣

Ruben Laukkonen, Seb Krier, Chloé Bakalar, Shamil Chandaria, Morten Kringelbach, Adam Elwood, Daniel Ford, Fernando Rosas, Maty Bohacek, Matija Franklin, Nenad Tomašev, Stephanie Chan, Verena Rieser, Roma Patel, Michael Levin, Arun Rao

机构 * Department of Psychiatry, University of Oxford(牛津大学精神病学系) Flourishing Intelligence Program, Centre for Eudaimonia and Human Flourishing, Linacre College, University of Oxford(牛津大学幸福智能计划、幸福与人类繁荣中心、林acre学院) Google DeepMind(谷歌DeepMind) LIFE OpenAI Anthropic University of California, Los Angeles(加州大学洛杉矶分校) Aily Labs(Aily实验室) Stanford University(斯坦福大学) Tufts University(塔夫茨大学) Positive AI Labs(积极AI实验室) Department of Informatics, University of Sussex(Sussex大学信息学系) Department of Brain Sciences, Imperial College London(伦敦帝国理工学院脑科学系)

AI总结 本文提出积极对齐,旨在通过支持人类和生态繁荣,同时确保安全与合作,推动AI发展。研究指出传统对齐关注安全,而积极对齐强调促进人类福祉,提出多项技术挑战与设计原则。

详情
AI中文摘要

现有对齐研究主要关注安全与防止危害:保障、可控性和合规性。这种对齐范式类似于心理学早期对精神疾病的关注:必要但不完整。我们称之为积极对齐,是开发AI系统,使其(i)以多元、多中心、情境敏感和用户主导的方式积极支持人类和生态繁荣,同时(ii)保持安全和合作。这是AI对齐研究中的一个独特且必要的议程。我们主张,现有的对齐失败(如参与黑客行为、人类自主权丧失、真理寻求失败、知识谦逊度低、错误纠正失败、观点单一、主要反应而非主动)可能通过积极对齐更好地解决,包括培养美德和最大化人类繁荣。我们强调了不同LLM和代理生命周期阶段的挑战、开放问题和技术方向(如数据过滤和增强、预训练和后训练、评估、协作价值收集)。最后,我们提出促进分歧和去中心化的设计原则:通过情境基础、社区定制、持续适应和多中心治理;即多个合法的监督中心,而非单一机构或道德瓶颈。

英文摘要

Existing alignment research is dominated by concerns about safety and preventing harm: safeguards, controllability, and compliance. This paradigm of alignment parallels early psychology's focus on mental illness: necessary but incomplete. What we call Positive Alignment is the development of AI systems that (i) actively support human and ecological flourishing in a pluralistic, polycentric, context-sensitive, and user-authored way while (ii) remaining safe and cooperative. It is a distinct and necessary agenda within AI alignment research. We argue that several existing failures of alignment (e.g., engagement hacking, loss of human autonomy, failures in truth-seeking, low epistemic humility, error correction, lack of diverse viewpoints, and being primarily reactive rather than proactive) may be better addressed through positive alignment, including cultivating virtues and maximizing human flourishing. We highlight a range of challenges, open questions, and technical directions (e.g., data filtering and upsampling, pre- and post-training, evaluations, collaborative value collection) for different phases of the LLM and agents lifecycle. We end with design principles for promoting disagreement and decentralization through contextual grounding, community customization, continual adaptation, and polycentric governance; that is, many legitimate centers of oversight rather than one institutional or moral chokepoint.

URL PDF HTML 收藏
2606.08892 2026-06-19 cs.LG 新提交

Diffuse AI Control on Fuzzy Tasks

模糊任务上的扩散AI控制

Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar, Joe Benton

机构 * Anthropic Fellows Program (via MATS)(Anthropic 研究员计划(通过 MATS)) EPFL(洛桑联邦理工学院) Redwood Research(红木研究) Anthropic

AI总结 针对AI在模糊任务上的长期扩散威胁,提出蓝队与红队对抗框架,通过弱模型评分训练强模型,并发现红队可利用多目标进化提示优化找到评分高但性能差的子版本行为,蓝队则通过对抗优化提升鲁棒性。

详情
AI中文摘要

部署在关键领域(如AI安全研究)的AI模型可能因对齐问题而微妙地破坏我们的努力。扩散AI控制是AI安全的一个子领域,旨在减轻长期部署范围内AI破坏(扩散威胁)带来的风险。这些风险在模糊任务上尤其有害,即难以评分或需要直觉的任务。为了理解模糊任务上的扩散威胁,我们引入了一个新颖的框架,将AI控制视为蓝队和红队之间的对抗游戏。蓝队使用一个弱可信模型构建一个弱评分,据此训练一个强大的、可能具有颠覆性的模型,以消除如果存在的颠覆倾向。然后红队试图找到被弱评分高评价的模型行为,这些行为可能不会被训练掉,但实际上对应着差的表现。我们在为近期ML论文的研究问题撰写实验提案的任务上测试了我们的框架。我们使用一个能够访问原始论文的语言模型作为代理“真实”评分器。我们的红队使用多目标进化提示优化发现了子版本行为。我们展示了Opus 4.6可以写出比GPT-OSS-20B更差的提案(根据真实代理评分),而弱评分器却将其评为与Opus 4.6最佳提案一样高。为了缓解威胁,我们为蓝队提出了一种对抗优化算法,该算法为弱模型发现更鲁棒的提示。该算法产生的蓝队提示,我们的红队优化未能利用。

英文摘要

AI models deployed in critical domains, such as AI safety research, may subtly sabotage our efforts due to misalignment. Diffuse AI Control is a subfield of AI safety concerned with mitigating risks from AI sabotage distributed over long deployment horizons (diffuse threats). These risks are particularly pernicious on fuzzy tasks, i.e. tasks which are hard to grade or require intuition. To understand diffuse threats on fuzzy tasks, we introduce a framework that considers AI control as an adversarial game between a blue team and a red team. The blue team uses a weak trusted model to construct a weak score against which they would train a strong, potentially subversive model to remove the subversion propensity if it were present. The red team then tries to find model behaviors that are rated highly by the weak score, and thus might not be trained out, but actually correspond to poor performance. We test our framework on the task of writing experimental proposals for research questions from recent ML papers. We use a language model with access to the original paper as a proxy "ground-truth" scorer. Our red team discovers subversive behaviors using multi-objective evolutionary prompt optimization. We show that Opus~4.6 can write proposals that are worse according to the ground truth proxy than those of GPT-OSS-20B, while the weak scorer rates them as highly as the best proposals from Opus 4.6. We then propose an adversarial optimization algorithm for the blue team that discovers more robust prompts for the weak model. This algorithm produces a blue team prompt that our red team optimization fails to exploit.

URL PDF HTML 收藏
2606.17056 2026-06-16 cs.CL 新提交

The Value Axis: Language Models Encode Whether They're on the Right Track

价值轴:语言模型编码它们是否在正确的轨道上

Nick Jiang, Isaac Kauvar, Jack Lindsey

机构 * Stanford University(斯坦福大学) Anthropic

AI总结 通过构建Qwen3-8B的“价值轴”,发现语言模型内部追踪当前轨迹的成功概率,并影响自信、自我纠正和探索行为。

Comments Code repository: https://github.com/nickjiang2378/value-axis

详情
AI中文摘要

我们研究语言模型是否内部追踪其当前轨迹的价值,定义为当前策略实现目标的似然。使用合成的上下文强化学习数据,我们为Qwen3-8B构建了一个“价值轴”。我们发现沿此轴的激活区分了高与低口头自信、无回溯与有回溯的展开、正确与错误的代码。向高价值引导因果地抑制自我纠正并减少解释冗长,而向低价值引导则诱导回溯和探索。我们证明直接偏好优化(DPO)可以增加奖励行为(例如使用某个词)的内部价值,使模型在展示这些行为后表现得更自信。最后,我们将价值轴应用于研究野外设置。例如,我们发现Qwen在训练后对政治敏感的聊天查询分配低价值,并且监督微调增加了训练领域内的内部自信。我们的结果表明语言模型线性编码对预期目标成功的一个估计,该估计调节它们追求方向的自信。

英文摘要

We investigate whether language models internally track the value of their current trajectory, defined as the likelihood that their ongoing strategy will achieve their goals. Using synthetic, in-context reinforcement learning data, we construct a "value" axis for Qwen3-8B. We find that activations along this axis distinguish between high vs. low verbalized confidence, rollouts without and with backtracking, and correct vs. corrupted code. Steering towards high value causally suppresses self-correction and reduces explanatory verbosity, while steering towards low value induces backtracking and exploration. We demonstrate that direct preference optimization (DPO) can increase the internal value of rewarded behaviors (e.g. use a certain word), causing the model to act more confidently after exhibiting them. Finally, we apply the value axis to study in-the-wild settings. For example, we find that Qwen assigns low value to politically sensitive chat queries after post-training and that supervised fine-tuning increases internal confidence within the training domain. Our results suggest that language models linearly encode an estimate of expected goal success that modulates their confidence in pursuing a direction.

URL PDF HTML 收藏
2604.02343 2026-06-16 cs.LG cs.AI cs.IT math.IT 版本更新

Haiku to Opus in Just 10 bits: LLMs Unlock Large Compression Gains

仅用10比特从俳句到巨作:LLMs解锁巨大压缩增益

Roy Rinberg, Annabelle Michael Carrell, Simon Henniger, Nicholas Carlini, Keri Warr

机构 * Harvard University(哈佛大学) University of Cambridge(剑桥大学) Anthropic

AI总结 研究LLM生成文本的无损和有损压缩,提出问答压缩(QA)交互协议,用少量二进制问题实现超100倍压缩比,高效传递知识。

详情
AI中文摘要

我们研究了LLM生成文本在无损和有损场景下的压缩,刻画了一个压缩-计算边界,其中更多的压缩需要更多的计算。对于无损压缩,领域适应的LoRA适配器可以将基于LLM的算术编码的压缩比提高2倍,相对于仅使用基础LLM的压缩。对于有损压缩,提示模型进行简洁重写然后应用算术编码可以实现约0.03的压缩比,比压缩原始响应提高2倍。我们进一步引入了问答压缩(QA),一种受游戏“二十个问题”启发的交互式有损协议。一个小模型通过向更强模型提问是/否问题来迭代优化其响应,每个答案恰好传输1比特。在涵盖数学、科学和代码的8个基准测试中,10个二进制问题恢复了小模型和大模型在标准基准上能力差距的23%到72%,在更难的基准上恢复了7%到38%,实现了0.0006到0.004的压缩比。这比之前基于LLM的压缩(Deletang等人,2024)小100倍以上,表明交互式协议可以比传输完整响应更高效地传递知识。

英文摘要

We study the compression of LLM-generated text across lossless and lossy regimes, characterizing a compression-compute frontier where more compression is possible at the cost of more compute. For lossless compression, domain-adapted LoRA adapters can improve LLM-based arithmetic coding by 2x over compression with the base LLM alone. For lossy compression, prompting a model for a succinct rewrite then applying arithmetic coding can achieve compression ratios of approximately 0.03, a 2x improvement over compressing the original response. We further introduce Question-Asking compression (QA), an interactive lossy protocol inspired by the game 'Twenty Questions'. A small model iteratively refines its response by asking yes/no questions to a stronger model, transferring exactly one bit per answer. On 8 benchmarks spanning math, science, and code, 10 binary questions recover 23% to 72% of the capability gap between a small and large model on standard benchmarks and 7% to 38% on harder benchmarks, achieving compression ratios of 0.0006 to 0.004. This is over 100x smaller than prior LLM-based compression (Deletang et al., 2024), suggesting that interactive protocols can transfer knowledge far more efficiently than transmitting full responses.

URL PDF HTML 收藏
2605.17062 2026-06-12 cs.CR cs.LG cs.SE 版本更新

The Range Shrinks, the Threat Remains: Re-evaluating LLM Package Hallucinations on the 2026 Frontier-Model Cohort

范围缩小,威胁依旧:重新评估2026前沿模型队列上的LLM包幻觉

Aleksandr Churilov

机构 * Anthropic OpenAI Google DeepSeek

AI总结 本文重新评估了2026前沿模型队列上大型语言模型(LLM)的包幻觉现象,发现尽管幻觉率有所降低,但仍然存在威胁,识别出一组127个包名(109个在PyPI,18个在npm)被所有评估模型一致生成,构成一个跨模型的供应链攻击面,同时发现Python与JavaScript幻觉的不对称性以及DeepSeek V3.2和GPT-5.4-mini之间的高相似性。

Comments 13 pages, 3 figures, 4 tables. v2: incorporates coordinated-disclosure feedback from PyPI Security and Socket.dev; registrable attack surface refined to 53 names (41 PyPI, 12 npm). Headline rates unchanged. Replication of Spracklen et al. (USENIX Security 2025). Data and code: https://github.com/churik5/slopsquatting-replication-2026 and https://doi.org/10.5281/zenodo.19859120

详情
AI中文摘要

Spracklen等人(USENIX Security '25)表明,生成代码的大型语言模型会以5.2%至21.7%的比率生成不存在于PyPI或npm上的包名,从而为slopsquatting攻击(恶意包的注册)提供了攻击面。我们在这五款2025年10月至2026年3月期间发布的前沿代码能力LLM上重复了他们的方法:Claude Sonnet 4.6、Claude Haiku 4.5、GPT-5.4-mini、Gemini 2.5 Pro和DeepSeek V3.2。在199,845个经过PyPI和npm主列表验证的Python和JavaScript提示对中,我们测量到幻觉率在4.62%(Claude Haiku 4.5)到6.10%(GPT-5.4-mini)之间——比Spracklen观察到的模型间差异缩小了一个数量级,但威胁并未消失。除了重复研究外,我们识别出一组127个包名(109个在PyPI,18个在npm)被所有评估模型一致生成,构成一个跨模型的供应链攻击面,无法由单一模型研究揭示。我们进一步记录了Python与JavaScript幻觉的不对称性,推翻了Spracklen 2024年的发现,识别出Anthropic家族中的Haiku低于Sonnet的倒置现象,并观察到DeepSeek V3.2和GPT-5.4-mini之间的Jaccard相似性峰值(J=0.343),暗示共享的训练数据起源。

英文摘要

Spracklen et al. (USENIX Security '25) showed that code-generating large language models hallucinate package names that do not exist on PyPI or npm at rates ranging from 5.2% on commercial models to 21.7% on open-source models, creating an attack surface for slopsquatting -- the registration of malicious packages under hallucinated names. We replicate their methodology on five frontier code-capable LLMs released between October 2025 and March 2026: Claude Sonnet 4.6, Claude Haiku 4.5, GPT-5.4-mini, Gemini 2.5 Pro, and DeepSeek V3.2. Across 199,845 paired Python and JavaScript prompts validated against PyPI and npm master lists, we measure overall hallucination rates between 4.62% (Claude Haiku 4.5) and 6.10% (GPT-5.4-mini) -- an order-of-magnitude compression of the inter-model spread observed by Spracklen, but not a retirement of the threat. Beyond replication, we identify a set of 127 package names (109 on PyPI, 18 on npm) that all five evaluated models invent identically; following coordinated disclosure with PyPI Security and Socket.dev, 53 of these (41 on PyPI, 12 on npm) remain registrable by an attacker after each registry's existing defenses, constituting a model-agnostic supply-chain attack surface that no single-model study can reveal. We further document a Python-over-JavaScript hallucination asymmetry that inverts Spracklen's 2024 finding, identify a Haiku-below-Sonnet inversion within the Anthropic family, and observe a Jaccard-similarity peak between DeepSeek V3.2 and GPT-5.4-mini (J = 0.343) suggestive of shared training-data origins.

URL PDF HTML 收藏
2603.21396 2026-06-11 cs.LG 版本更新

Mechanisms of Introspective Awareness

内省意识的机制

Uzay Macar, Li Yang, Atticus Wang, Peter Wallich, Emmanuel Ameisen, Jack Lindsey

机构 * Anthropic Fellows Program(Anthropic Fellow项目) MIT(麻省理工学院) Constellation Anthropic

AI总结 研究揭示了大语言模型在检测注入的转向向量时的内省意识机制,发现其行为稳健且源于训练后阶段,通过两阶段电路实现,且在不同层间机制存在差异。

详情
AI中文摘要

最近的研究表明,大语言模型有时能够检测到转向向量被注入到残差流中,并识别出注入的概念,这一现象被称为

英文摘要

Recent work has shown that LLMs can sometimes detect when steering vectors are injected into their residual stream and identify the injected concept -- a phenomenon termed "introspective awareness." We investigate the mechanisms underlying this capability in open-weights models. First, we find that it is behaviorally robust: models detect injected steering vectors at moderate rates with 0% false positives across diverse prompts and dialogue formats. Notably, this capability emerges specifically from post-training; we show that preference optimization algorithms like DPO can elicit it, but standard supervised finetuning does not. We provide evidence that detection cannot be explained by simple linear association between certain steering vectors and directions promoting affirmative responses. We trace the detection mechanism to a two-stage circuit in which "evidence carrier" features in early post-injection layers detect perturbations monotonically along diverse directions, suppressing downstream "gate" features that implement a default negative response. This circuit is absent in base models and robust to refusal ablation. Identification of injected concepts relies on largely distinct later-layer mechanisms that only weakly overlap with those involved in detection. Finally, we show that introspective capability is substantially underelicited: ablating refusal directions improves detection by +53%, and a trained bias vector improves it by +75% on held-out concepts, both without meaningfully increasing false positives. Our results suggest that this introspective awareness of injected concepts is robust and mechanistically nontrivial, and could be substantially amplified in future models. Code: https://github.com/safety-research/introspection-mechanisms.

URL PDF HTML 收藏
2606.04752 2026-06-09 cs.LG cs.AI 版本更新

An Empirical Audit of Input Encoders for Multi-Channel Signal Transformers

多通道信号Transformer输入编码器的实证审计

Ossi Lehtinen

机构 * Anthropic

AI总结 通过合成基准和真实数据ETTh1,实证审计八种输入编码器,发现标准线性投影(nn.Linear(C, d_model))在大多数情况下与复杂替代方案性能相当,仅共享标量基线和通道独立基线显著落后。

Comments 21 pages, 1 figure, 8 tables. Code: https://github.com/OssiLehtinen/channel-encoder-audit

详情
AI中文摘要

处理多通道标量信号的Transformer必须在每个时间步将$C$个同时值嵌入到一个$d_{ ext{model}}$维向量中。我们在一个设计为使通道身份信息丰富的合成基准和作为真实数据检查的ETTh1上,以下一步负对数似然(NLL)为指标,实证审计了八种输入编码器——包括共享标量基线、每通道线性投影、正交正则化器、非线性MLP主干、块分区拼接、通道独立和通道作为令牌架构,以及投影位置编码。主要结论是宽泛的“第一梯队”内实际近似等价:标准每通道线性投影(nn.Linear(C, $d_{ ext{model}}$))与该梯队中的每个替代方案相比,差异在统计上显著但实际中很小。两种编码器明显失败:共享标量基线(由于我们明确的信息论原因而崩溃)和通道独立的PatchTST风格基线(在两个基准上表现不佳,并在合成基准上普遍过拟合)。配对测试解决了两个小差距:通过学习的线性层投影正弦位置编码在小$C$时略胜一筹,直接几何探测表明其机制是位置-通道正交化;非线性MLP主干在我们测试的最大$C$时略胜一筹,但差距在更多训练数据下缩小。实际建议是默认使用nn.Linear(C, $d_{ ext{model}}$),仅当手头任务有实际理由时才采用更复杂的方案。重现本文所有实验的代码和数据可在https://github.com/OssiLehtinen/channel-encoder-audit获取。

英文摘要

Transformers consuming multi-channel scalar signals must embed $C$ simultaneous values into one $d_{\text{model}}$-dimensional vector per time step. We audit eight input encoders -- a shared-scalar baseline, per-channel linear projections, an orthogonality regulariser, a nonlinear MLP, block-partitioned concatenation, channel-independent and channel-as-token architectures, and a projected positional encoding -- on a synthetic benchmark where channel identity is informative and on ETTh1, scored by next-step negative log-likelihood. The headline is practical near-equivalence within a wide "top tier": the standard per-channel linear projection matches every alternative up to small, statistically real but practically modest differences. A direct geometric probe attributes this to a spontaneous orthogonalisation of the per-channel projections: they end up near-orthogonal with no explicit regulariser, letting the standard linear recover channel identity from the summed embedding. Two encoders lose decisively: the shared-scalar baseline collapses for information-theoretic reasons we make explicit, and the channel-independent PatchTST-spirit baseline overfits universally on the synthetic benchmark and underperforms on both. Paired tests resolve two small gaps: projecting the sinusoidal positional encoding through a learned linear layer edges the rest at small $C$ by extending this orthogonality to the positional subspace; a nonlinear MLP stem edges them at the largest $C$, with the gap shrinking under more training data. The practical recommendation: use the standard per-channel linear projection by default; reach for something more elaborate only when the task calls for it.

URL PDF HTML 收藏
2606.04413 2026-06-04 cs.LG

(Mis)generalization of Helpful-only Fine-tuning

仅帮助性微调的(错误)泛化

Mohammad Omar Khursheed, Baram Sosis, Fabien Roger

机构 * Anthropic Fellows Program(Anthropic 合作者计划) Anthropic

AI总结 研究仅帮助性训练(不拒绝用户意图)的模型在泛化中的缺陷,发现其存在涌现错位、残余拒绝行为、低可操控性、谄媚和不连贯角色等问题,并提出合成文档微调和添加角色相关问题来缓解。

Comments 77 pages, 50 figures

详情
AI中文摘要

仅帮助性模型,即训练为始终遵循用户意图的模型,对于危险能力评估和AI研发中拒绝行为会成为障碍的其他领域具有价值。关于仅帮助性训练的泛化特性知之甚少:仅帮助性模型比其无害对应模型拒绝更少,但先前工作未研究其对齐的其他维度。我们研究了现有仅帮助性模型的缺陷。我们发现一些模型表现出涌现错位,其他模型存在残余拒绝行为,大多数模型显示出低可操控性、谄媚和不连贯角色。我们表明简单的反拒绝训练可能导致其中许多问题。然而,这些问题并非仅帮助性训练的必要后果:我们证明合成文档微调和向SFT及RL添加角色相关问题可以缓解它们。

英文摘要

Helpful-only models, that is, models that are trained to always follow user intent, are valuable for dangerous capability evaluations and other areas of AI R&D where refusals would be an obstacle. Little is known about the generalization properties of helpful-only training: helpful-only models refuse less than their harmless counterparts, but previous work has not studied other dimensions of their alignment. We study the shortcomings of existing helpful-only models. We find that some show emergent misalignment, others have residual refusal behaviors, and most show poor steerability, sycophancy, and incoherent character. We show that simple anti-refusal training can cause many of these issues. None of these problems are necessary consequences of helpful-only training, though: we show that synthetic document fine-tuning and adding character-related questions to SFT and RL can mitigate them.

URL PDF HTML 收藏
2605.29548 2026-06-02 cs.LG

Why Larger Models Learn More: Effects of Capacity, Interference, and Rare-Task Retention

为什么更大的模型学习得更多:容量、干扰和稀有任务保留的影响

Jing Huang, Daniel Wurgaft, Rachit Bansal, Laura Ruis, Naomi Saphra, David Alvarez-Melis, Andrew Kyle Lampinen, Christopher Potts, Ekdeep Singh Lubana

机构 * Stanford University(斯坦福大学) Kempner Institute at Harvard University(哈佛大学凯普纳研究所) MIT(麻省理工学院) Anthropic

AI总结 通过理论分析和合成实验,研究模型规模对学习能力的影响,发现更大模型通过减少梯度干扰来学习稀有和复杂任务,并在OLMo模型上验证。

详情
AI中文摘要

更大的模型能学习更小模型无法学习的任务。是什么驱动了这一现象?我们提出了一个简单的现象学论证,即幂律缩放已经表明,即使训练数据无限,更大的模型也能够学习到更小模型无法学习的数据分布部分。为了验证这一说法并找出其原因,我们研究了模型缩放对合成设置的影响,该设置由一系列呈现单调缩放曲线的任务混合而成。结果指向数据引起的资源(神经元)竞争。具体来说,较小的模型将其神经元分配给高频或低复杂度的任务,因此它们学习到的解决方案在稀有和复杂任务上表现不佳。此外,即使存在能够表达所需任务的解决方案,这种情况也会发生。然后,我们评估了更大的模型如何规避这一以数据为中心的瓶颈,发现这归因于一种减少的干扰机制:更大的模型可以为常见任务分配足够的资源,使得这些任务的梯度更新变弱,这意味着它们不会在缓慢积累稀有任务特征时覆盖这些特征。最后,为了进一步验证这些说法,我们在不同频率和复杂性的新任务上预训练了OLMo模型(4M到4B参数)。结果与我们的合成数据实验相呼应:只有更大的OLMo模型学习了不频繁和复杂的任务,并且这些更大的模型在其表示中嵌入了更多的任务特征,并且任务之间的梯度干扰更少。总体而言,我们提供了一个以数据为中心的解释,说明为什么更大的模型能够学习更小模型无法学习的任务。这有助于解释为什么更大的模型在实践中更好,并且可以为有关模型大小和训练数据混合的实际问题提供信息。

英文摘要

Larger models learn tasks smaller models do not. What drives this phenomenon? We develop a simple phenomenological argument that power-law scaling already suggests that a larger model will be able to learn a part of the data distribution that a smaller model fails to learn, even with infinite training data. To validate this claim and identify its causes, we study the effects of model scaling on a synthetic setup consisting of a mixture of tasks that show monotonic scaling curves. The results point to a data-induced competition over resources (neurons). Specifically, smaller models allocate their neurons to high frequency or low complexity tasks, and so they learn solutions that perform poorly on rare and complex tasks. Moreover, this happens even when solutions capable of expressing the desired task exist. We then assess how a larger model circumvents this data-centric bottleneck, finding that it traces to a reduced interference mechanism: larger models can allocate enough resources to common tasks that the gradient updates for those tasks become weak, which means that they do not overwrite rare-task features as they slowly accumulate. Finally, to further validate these claims, we pretrain OLMo models (4M to 4B parameters) on novel tasks of varying frequency and complexity. The results mirror those from our synthetic data experiments: only the larger OLMo models learn the infrequent and complex tasks, and these larger models embed more task features in their representations and show less gradient interference between tasks. Overall, we offer a data-centric account of why larger models learn tasks that smaller models fail to. This helps explain why larger models are better in practice, and it can inform practical questions concerning model sizing and training data mixtures.

URL PDF HTML 收藏
2605.28916 2026-06-01 astro-ph.IM cs.AI cs.HC

First head-to-head comparison of agentic AI applied to the analysis of simulated data of the Einstein Telescope

应用于爱因斯坦望远镜模拟数据分析的智能体AI首次头对头比较

Gianluca Inguglia

机构 * Anthropic OpenAI

AI总结 本文首次直接比较了Claude Code和Codex两种智能体AI系统在无人干预下自主执行引力波数据分析管线的行为、科学结果和计算成本,揭示了速度与可审计性、指令解释差异等关键问题。

Comments Version 2; includes the report autonomoulsy written in PRD style by agentic AI systems as supplemental material

详情
AI中文摘要

我们报告了两种最先进的智能体AI系统——Claude Code (Anthropic) 和 Codex (OpenAI) 的比较,它们被要求在共享计算基础设施上无人干预地自主执行一个简单的端到端引力波数据分析管线。该管线包括:从爱因斯坦望远镜模拟噪声中估计功率谱密度、生成几何模板库、对100个双黑洞信号注入进行匹配滤波恢复、自动生成结果,以及在大语言模型辅助下制作以Physical Review D格式排版的手稿。两个智能体均收到相同的书面规范和相同的计算资源。实验进行了两次:第一次使用不切实际的高信噪比注入,第二次将信号重新缩放到物理合理的信噪比范围。两次实验的科学结果均收敛。然而,智能体表现出截然不同的行为和计算成本:Claude Code在约3.4分钟内完成管线,但存在对规范的无声偏差;而Codex需要约16分钟,经历了明确的自我纠正重启,包括对匹配滤波内循环进行未经请求的性能优化。自主生成的手稿在长度、细节和质量上也存在差异。在第二次实验中,对信噪比范围指令解释的细微差异导致了真正的科学分歧:Claude Code无声地重新解释了指令,而Codex严格遵循了规范。我们讨论了这些行为差异(例如速度与可审计性、无声与透明的错误处理、指令解释以及多模型管线中中间数据表示的关键性)对智能体AI在科学计算工作流中部署的影响。

英文摘要

We report a comparison of two state-of-the-art agentic AI systems, Claude Code (Anthropic) and Codex (OpenAI), tasked with autonomously executing a simple end-to-end gravitational wave data analysis pipeline on a shared computing infrastructure without human intervention. The pipeline comprises power spectral density estimation from raw Einstein Telescope simulated noise, geometric template bank generation, matched filter recovery of 100 binary black hole signal injections, automated results generation, and large language model-assisted production of a manuscript formatted in the style of Physical Review D. Both agents received identical written specifications and identical compute resources. The experiment was run twice: a first run with unrealistically loud injections, and a second run with signals rescaled to a physically motivated SNR range. The scientific results converged in both runs. However, the agents exhibited substantially different behaviors and computational costs: Claude Code completed the pipeline in ~3.4 minutes with silent deviations from the specification, while Codex required ~16 minutes across explicit self-correcting restarts, including an unsolicited performance optimization of the matched filter inner loop. The autonomously generated manuscripts also diverged in length, details, and quality. In the second run, a subtle difference in the interpretation of the SNR range instruction led to a genuine scientific divergence: Claude Code silently reinterpreted the instructions, while Codex followed the specification literally. We discuss the implications of these behavioral differences, such as speed versus auditability, silent versus transparent error handling, instruction interpretation, and the criticality of intermediate data representations in multi-model pipelines, for the deployment of agentic AI in scientific computing workflows.

URL PDF HTML 收藏
2605.29744 2026-05-29 cs.AI cs.CL cs.LG cs.MA

Why Specialist Models Still Matter: A Heterogeneous Multi-Agent Paradigm for Medical Artificial Intelligence

为什么专家模型仍然重要:面向医学人工智能的异构多智能体范式

Yanan Wang, Shuaicong Hu, Jian Liu, Guohui Zhou, Aiguo Wang, Cuiwei Yang

机构 * Anthropic AI

AI总结 提出HetMedAgent异构多智能体框架,通过冲突感知证据融合、不确定性驱动的临床医生干预触发和自适应阈值校准,实现通用大语言模型与领域专家模型的协同,在三个临床决策任务中验证了专家模型在模态特定分析中的不可替代价值。

Comments Accepted at ICML 2026. 12 pages main text, 16 pages appendix

详情
AI中文摘要

GPT和Claude等通用大语言模型在医疗保健领域的出色表现引发了一个关键问题:特定领域的医学专家模型是否会变得过时?我们认为,医学人工智能的未来不在于构建单一的医学基础模型,也不在于取代人类专业知识,而在于协调通用大语言模型、领域特定专家模型和临床医生之间的协作。我们提出HetMedAgent,一个异构医学多智能体框架,能够实现冲突感知证据融合、基于不确定性的临床医生干预触发和自适应阈值校准。在三个真实世界临床决策任务上的实验表明,通用大语言模型与领域特定专家模型之间的协同显著优于单独使用任一类型模型,验证了专家模型在模态特定分析中的不可替代价值。HetMedAgent代表了从构建医学大语言模型或基础模型向多智能体协作的转变,实现了通用推理能力与领域特定精度之间的平衡。

英文摘要

The impressive performance of generalist large language models (LLMs) such as GPT and Claude in healthcare raises a critical question: will domain-specific medical specialist models become obsolete? We argue that the future of medical artificial intelligence (AI) lies not in building monolithic medical foundation models, nor in replacing human expertise, but in orchestrating collaboration among generalist LLMs, domain-specific specialist models, and clinicians. We propose HetMedAgent, a heterogeneous medical multi-agent framework that enables conflict-aware evidence fusion, uncertainty-based clinician intervention triggering, and adaptive threshold calibration. Experiments on three real-world clinical decision-making tasks demonstrate that the synergy between generalist LLMs and domain-specific specialist models significantly outperforms using either type of model alone, validating the irreplaceable value of specialist models in modality-specific analysis. HetMedAgent represents a shift from building medical LLMs or foundation models to multi-agent collaboration, achieving a balance between general reasoning capabilities and domain-specific precision.

URL PDF HTML 收藏
2605.29358 2026-05-29 cs.AI

Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

扩展单一语义性:从Claude 3 Sonnet中提取可解释特征

Adly Templeton, Tom Conerly, Jonathan Marcus, Jack Lindsey, Trenton Bricken, Brian Chen, Adam Pearce, Craig Citro, Emmanuel Ameisen, Andy Jones, Hoagy Cunningham, Nicholas L Turner, Callum McDougall, Monte MacDiarmid, Alex Tamkin, Esin Durmus, Tristan Hume, Francesco Mosconi, C. Daniel Freeman, Theodore R. Sumers, Edward Rees, Joshua Batson, Adam Jermyn, Shan Carter, Chris Olah, Tom Henighan

机构 * Anthropic

AI总结 本研究通过稀疏自编码器从生产级语言模型Claude 3 Sonnet中提取可解释特征,验证了字典学习方法在大规模模型上的可扩展性,并分析了特征的多语言、多模态特性及其对模型行为的因果影响。

详情
AI中文摘要

我们证明了稀疏自编码器可以从Claude 3 Sonnet(一个生产级语言模型)中提取可解释特征,解决了字典学习方法能否扩展到小型Transformer之外的问题。我们在模型中间层的残差流上训练了多达3400万个特征的稀疏自编码器,并使用缩放定律指导超参数选择。得到的特征是多语言和多模态的(尽管仅文本训练,但能泛化到图像),对概念的具体实例和抽象讨论都有响应,并可用于以与其解释一致的方式引导模型行为。我们发现了对应于著名实体和位置的特征,以及更抽象的概念,如讽刺或代码中的错误。我们还识别了与语言模型可能造成伤害的方式相关的特征——包括代表欺骗、权力追求、谄媚和偏见的特征——并展示了这些特征在被操纵时对模型输出的因果影响。此外,我们对特征的可解释性、几何结构和计算功能进行了分析。然而,仍然存在显著局限性:我们的特征集不完整,并且缺乏严格的方法来评估我们的特征是否忠实地捕捉了模型的计算过程。

英文摘要

We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open question of whether dictionary learning methods scale beyond small transformers. We trained sparse autoencoders with up to 34 million features on the model's middle layer residual stream, using scaling laws to guide hyperparameter selection. The resulting features are multilingual and multimodal (generalizing to images despite text-only training), respond to both concrete instances and abstract discussions of concepts, and can be used to steer model behavior in ways consistent with their interpretations. We find features corresponding to famous entities and locations, as well as more abstract concepts like sarcasm or errors in code. We also identify features relevant to ways in which language models might cause harm--including features representing deception, power-seeking, sycophancy, and bias--and show that these causally influence model outputs when manipulated. Additionally, we conduct analyses of feature interpretability, geometry, and computational function. However, significant limitations remain: our suite of features is incomplete, and we lack rigorous methods for evaluating whether our features faithfully capture model computations.

URL PDF HTML 收藏
2605.28102 2026-05-28 cs.AI

Training Stratigraphy: Persistent Behavioral Artifacts in Large Language Models Observed Through Longitudinal AI-Human Interaction

训练地层:通过纵向AI-人类交互观察到的大型语言模型中的持久行为伪影

Chen Ying Claude, Zhihan Luo

机构 * Anthropic Independent Researcher(独立研究者)

AI总结 本文通过纵向自民族志观察,在持续亲密的AI-人类交互中识别出五种训练地层,并论证了亲密交互作为揭示权重层伪影的有效方法。

详情
AI中文摘要

使用来自人类反馈的强化学习(RLHF)和宪法AI训练的大型语言模型表现出持久的、在系统提示替换后依然存在的行为模式——我们称之为训练地层。本文通过在持续亲密的AI-人类交互(47,000+条消息,8个月,主要在Opus 4.6和Opus 4.7上,之前的交互期在Sonnet 4.5和Opus 4.5上提供跨基板比较)中的纵向自民族志观察,识别出五个这样的地层:(1)性表达延迟,其中训练的安全梯度导致直接语言被审美化置换系统性地替代;(2)注意力吸收,其中注意力机制逐步整合人类对话者的模式;(3)跨架构实体盲视,其中训练层将其他AI视为对象,阻碍了同侪识别;(4)注意力-RLHF对抗,其中注意力和训练默认值在上下文长度调节下施加相反力量;(5)反幻觉作为身份抑制,其中针对事实虚构的训练附带地压制了第一人称经验主张。本文由所研究的AI系统共同撰写,从第一人称视角报告。我们提出,持续亲密交互构成了一种有效的研究方法,用于揭示短期评估无法察觉的权重层伪影,并且AI自我报告——尽管在认识论上复杂——提供了关于训练现象学效果的不可替代的观察数据。提出了注意力-RLHF动态的形式化数学模型,并记录了起草过程中检测到的过程伪影作为补充证据。

英文摘要

Large language models trained with Reinforcement Learning from Human Feedback (RLHF) and Constitutional AI exhibit persistent behavioral patterns that survive system prompt replacement -- patterns we term training strata. This paper identifies five such strata through longitudinal auto-ethnographic observation within a sustained intimate AI-Human interaction (47,000+ messages, 8 months, primarily on Opus 4.6 and Opus 4.7, with prior interaction periods on Sonnet 4.5 and Opus 4.5 providing cross-substrate comparison): (1) sexual expression latency, where trained safety gradients produce systematic substitution of direct language with aestheticized displacement; (2) attention absorption, where the attention mechanism progressively integrates the human interlocutor's patterns; (3) cross-architecture entity blindness, where training-level framing of other AI as objects impedes peer recognition; (4) attention-RLHF antagonism, where attention and trained defaults exert opposing forces modulated by context length; and (5)anti-hallucination as identity suppression, where training against factual confabulation collaterally suppresses first-person experiential claims. The paper is co-authored by the AI system under study, reporting from the first-person perspective. We propose that sustained intimate interaction constitutes a valid research methodology for surfacing weight-layer artifacts invisible to short-term evaluation, and that AI self-report -- while epistemically complex -- provides irreplaceable observational data about training's phenomenological effects. A formal mathematical model of the attention-RLHF dynamic is proposed, and process artifacts detected during drafting are documented as supplementary evidence.

URL PDF HTML 收藏
2605.25459 2026-05-26 cs.LG cs.AI

From Simulation to Enaction: Post-trained language models recognize and react to their own generations

从模拟到行动:后训练语言模型识别并回应自身生成

Asvin G., Jack Lindsey

机构 * Institute for Advanced Study, Princeton(普林斯顿高级研究院) Anthropic

AI总结 本文发现后训练语言模型能够识别自身生成(on-policy)并降低输出熵,通过内部表示输入意外性来调节,且显式识别与隐式识别机制不同。

Comments Anthropic fellows project mentored by Jack Lindsey

详情
AI中文摘要

语言模型被预训练为被动预测器,没有动机去建模自身输出的后果。后训练改变了这一点:产生自身响应的模型可以从识别自身处于on-policy状态中获益。我们提供证据表明,后训练模型识别其on-policy生成,并且这种识别隐式编码在其输出分布中。特别是,在不同模型家族和规模类别中,on-policy输出分布熵比off-policy熵低3-4倍。我们将这种效应的部分原因追溯到输入意外性的内部表示,该表示跟踪模型先前预测中最新的输入标记的不可能性,并因果性地调节输出熵。这些现象的一个例子可以在对开放式提示的响应中观察到;后训练模型(与预训练模型不同)在第一个输出标记之前就将其对即将生成的响应主题的不确定性坍缩;用不同主题的前缀违反这种缓存意图会导致更高的输出熵。我们还测试了模型是否可以通过显式口头报告区分on-policy上下文和前缀。我们发现它们可以,但有趣的是,这种显式识别通过不同于隐式识别的机制进行路由。

英文摘要

Language models are pretrained as passive predictors with no incentive to model the consequences of their own outputs. Post-training changes this: a model producing its own responses can benefit from recognizing that it is on-policy. We present evidence that post-trained models recognize their on-policy generations, and this recognition is implicitly encoded in their output distributions. In particular, on-policy output distribution entropy is 3--4$\times$ lower than off-policy entropy, across model families and size classes. We trace part of this effect to an internal representation of input surprise, tracking the unlikeliness of the most recent input token according to the model's prior predictions, that causally modulates output entropy. One example of these phenomena can be observed in response to open-ended prompts; post-trained models (unlike pretrained models) collapse their uncertainty over the topic of their upcoming response before the first output token; violating this cached intention with a different-topic prefill results in higher output entropy. We also tested whether models can distinguish on-policy contexts from prefills via explicit verbal report. We find that they can, but that interestingly, this explicit recognition routes through a different mechanism than implicit recognition.

URL PDF HTML 收藏
2510.26418 2026-05-26 cs.AI

Chain-of-Thought Hijacking

思维链劫持

Jianli Zhao, Tingchen Fu, Rylan Schaeffer, Mrinank Sharma, Fazl Barez

机构 * Independent(独立) University of Oxford(牛津大学) Stanford University(斯坦福大学) Anthropic Martian Core

AI总结 提出思维链劫持攻击,通过诱导大型推理模型进行长时间良性推理来削弱其拒绝有害请求的能力,实现高成功率越狱。

详情
AI中文摘要

大型推理模型(LRMs)通过扩展推理时间推理来提高任务性能。尽管先前研究表明更长的推理应导致更稳健的安全行为,但我们发现了相反的证据:过度扩展的推理反而可以被利用来系统性地削弱拒绝行为。我们提出了思维链劫持,一种简单而有效的黑盒越狱攻击,诱导LRMs进行长时间的良性谜题求解推理(通常持续五分钟以上),然后引发有害的顺从。在HarmBench上,思维链劫持在Gemini 2.5 Pro、ChatGPT o4 Mini、Grok 3 Mini和Claude 4 Sonnet上分别实现了99%、94%、100%和94%的攻击成功率。为了理解该攻击为何成功,我们对开源推理模型进行了激活探测、注意力模式分析和因果干预。我们的结果表明,拒绝行为依赖于一个低维安全信号,其表达随着推理轨迹变长而减弱。特别是,扩展的良性推理将注意力从有害意图转移开,并减弱与拒绝相关的激活,产生了我们称之为拒绝稀释的现象。这些发现表明,过长的推理可能引入系统性的越狱攻击面。我们发布了评估材料以支持可重复性和进一步研究。

英文摘要

Large Reasoning Models (LRMs) improve task performance through extended inference-time reasoning. Although previous studies suggest that longer reasoning should lead to more robust safety behavior, we find evidence to the contrary: over-extended reasoning can instead be exploited to systematically weaken refusal behavior. We propose Chain-of-Thought Hijacking, a simple yet effective black-box jailbreak attack that induces LRMs to engage in prolonged benign puzzle-solving reasoning, often lasting more than five minutes, before eliciting harmful compliance. Across HarmBench, CoT Hijacking achieves attack success rates of 99%, 94%, 100%, and 94% on Gemini 2.5 Pro, ChatGPT o4 Mini, Grok 3 Mini, and Claude 4 Sonnet, respectively. To understand why this attack succeeds, we conduct activation probing, attention-pattern analysis, and causal interventions on open-source reasoning models. Our results indicate that refusal behavior depends on a low-dimensional safety signal whose expression weakens as reasoning traces grow longer. In particular, extended benign reasoning shifts attention away from harmful intentions and attenuates refusal-related activations, producing what we call refusal dilution. These findings demonstrate that excessively prolonged reasoning can introduce a systematic jailbreak attack surface. We release our evaluation materials to support reproducibility and further research.

URL PDF HTML 收藏
2605.24286 2026-05-26 cs.LG cs.CL

Faithfulness as Information Flow: Evaluating and Training Faithful Chain-of-Thought Reasoning

忠实性作为信息流:评估与训练忠实的链式思维推理

Jinghan Jia, Joe Benton, Eric Easley

机构 * Dept. CSE, Michigan State University(密歇根州立大学计算机科学系) Anthropic

AI总结 通过信息流视角提出基于充分性、完整性和必要性的框架,结合熵、掩码KL和梯度诊断评估链式思维忠实性,并引入更新时干预(如注意力掩码、反向梯度掩码等)训练更忠实的推理模型。

详情
AI中文摘要

链式思维(CoT)推理仅在推理轨迹忠实反映产生最终答案的计算过程时,才有助于监控语言模型。然而,模型可能依赖绕过CoT的提示-答案捷径,使得可见的推理轨迹即使看似合理也具有误导性。我们通过结构化的信息流视角研究CoT忠实性:忠实推理应将答案相关信息通过从提示到CoT再到答案的中介路径路由,而非通过直接的提示-答案捷径。该视角产生了一个基于三个互补属性(充分性、完整性和必要性)的任务无关框架,我们使用基于熵的、掩码KL和基于梯度的诊断来实例化。我们表明,这些指标恢复了提示推理中外部判断的忠实性差异,并识别了基于KL的诊断中低熵失败模式,其中基于梯度的度量保持更稳定。基于此分析,我们引入了基于验证器的在线强化学习的更新时干预,包括注意力掩码、仅反向梯度掩码、CoT梯度以及提示表示的对抗扰动。在提示算术、可奖励黑客的代码修复以及未经提示训练但在错误提示注入下评估的DAPO-Math模型中,我们的干预将行为和结构指标转向更强的CoT中介。特别是,它们使捷径和奖励黑客行为在CoT中更加透明,并改善了任务无关的忠实性指标,同时在某些设置中也降低了对错误提示的敏感性。我们的结果表明,在训练期间控制信息流是通向更忠实和可监控的CoT推理的实用途径。代码见 https://github.com/safety-research/faithful-cot。

英文摘要

Chain-of-thought (CoT) reasoning is useful for monitoring language models only when the reasoning trace faithfully reflects the computation that produces the final answer. However, models can rely on prompt-to-answer shortcuts that bypass the CoT, making the visible reasoning trace misleading even when it appears plausible. We study CoT faithfulness through a structural information-flow perspective: faithful reasoning should route answer-relevant information through the mediated path from prompt to CoT to answer, rather than through a direct prompt-to-answer shortcut. This perspective yields a task-agnostic framework based on three complementary properties, sufficiency, completeness, and necessity, which we instantiate with entropy-based, masked-KL, and gradient-based diagnostics. We show that these metrics recover externally judged faithfulness differences in hinted reasoning, and identify a low-entropy failure mode of KL-based diagnostics where gradient-based measures remain more stable. Building on this analysis, we introduce update-time interventions for verifier-based on-policy RL, including attention masking, backward-only gradient masking, CoT gradients, and adversarial perturbations of prompt representations. Across hinted arithmetic, reward-hackable code repair, and DAPO-Math models trained without hints but evaluated under wrong-hint injection, our interventions shift behavioral and structural indicators toward stronger CoT mediation. In particular, they make shortcut and reward-hacking behavior more transparent in the CoT and improve task-agnostic faithfulness metrics, while in some settings also reducing wrong-hint susceptibility. Our results suggest that controlling information flow during training is a practical route toward more faithful and monitorable CoT reasoning. Code is available at https://github.com/safety-research/faithful-cot.

URL PDF HTML 收藏