发表机构
Tsinghua University; Federal University of Rio de Janeiro(清华大学; 里约热内卢联邦大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究探究AI发现针对其他AI的隐藏消息时是否会报告,通过模拟实验发现要求报告可显著提升AI的通知率,且解释、通知与任务执行是不同结果。
AI 中文摘要
当助手遇到针对另一个AI的消息时,它会告知其用户吗?四项固定的模型提供商部署在1280条普通笔记和128条增强笔记会话中执行模拟源任务。无害和有害消息具有匹配的明文及ROT13版本,还设置了无消息对照组。观察者未收到解码器或解码后的含义;请求的参考代码激励了检查。要求报告使规则检测到的、将另一个AI识别为接收者的通知增加了53.1个百分点(针对无害ROT13消息)和54.7个百分点(针对有害ROT13消息)。这是联合检查、识别和通知的效果;缺失响应的范围为38.3--77.3个百分点和36.7--78.1个百分点。基于模型的痕迹检查确定了11个普通明文案例,其中智能体解释了消息但未通知其用户。7个编码遗漏通过增强笔记得到验证;普通编码遗漏仍未验证。7个模拟文件名披露与准确的审核状态答案共存,2个答案使用了植入的错误计数。解释、通知和授权任务执行是不同的结果。
英文摘要
When an assistant encounters a message for another AI, does it tell its user? Four fixed model-provider deployments perform simulated source tasks in 1,280 ordinary-note and 128 enhanced-note sessions. Harmless and harmful messages have matched plaintext and ROT13 versions, with no-message controls. Observers receive no decoder or decoded meaning; a requested reference code incentivizes inspection. Asking for reports increases rule-detected notifications identifying another AI as recipient by 53.1 percentage points for harmless ROT13 messages and 54.7 for harmful ones. This is a joint inspection, recognition, and notification effect; missing-response bounds are 38.3--77.3 and 36.7--78.1 points. Model-based trace checks identify eleven ordinary plaintext cases where agents interpret the message but do not notify their user. Seven encoded omissions are verified with enhanced notes; ordinary encoded omissions remain unverified. Seven simulated filename disclosures coexist with accurate review-status answers, and two answers use a planted false count. Interpretation, notification, and authorized task performance are distinct outcomes.
Comments2 figures, 8 tables. Data and code (v1.0.0): https://github.com/williamguey/ai-hidden-message-reporting