AI 中文总结
本研究提出合作性信号传递游戏“仅供您查看”,评估隔离模型实例间的协调能力,发现多数模型在避免可检测信号时协调困难,而前沿模型仍保持高性能,且跨架构协调弱于架构内。
AI 中文摘要
随着模型生成的内容在自动化工作流中越来越多地被其他模型实例消费,一个具有实际重要性的问题应运而生:一个模型能否在自然语言中嵌入一种信号,使得同一模型的独立实例仅依靠共享的预训练和任务指令即可检测到该信号,而无需任何共享记忆或针对协调性的专门训练?我们提出了“仅供您查看”(For Your Eyes Only),一个旨在直接评估这一能力的合作性信号传递游戏。发送者(Sender)为两个单词生成自由形式的描述,其中一个是隐藏目标;一个隔离的接收者(Receiver)必须识别出该目标。我们使用来自既定心理语言学语料库的300个词对,评估了来自四个架构家族的七个当代模型,并采用双遍成功率(Double-Pass Success Rate)来控制输出偏差。我们发现,一旦要求模型避免可检测的信号,大多数模型难以维持协调性,而一个前沿模型即使在经过此类过滤后仍能保持近乎完美的性能。我们进一步表明,模型可以将这种能力导向故意的误导,并且跨架构的协调性始终弱于架构内部的协调性。
英文摘要
As model-generated content is increasingly consumed by other model instances in automated workflows, a practically important question arises: can a model embed a signal in natural language that an independent instance of the same model can detect, relying only on shared pre-training and task instructions, without any shared memory or coordination-specific training? We introduce For Your Eyes Only, a cooperative signalling game designed to evaluate this directly. A Sender produces free-form descriptions for two words, one of which is a hidden target; an isolated Receiver must identify it. We evaluate seven contemporary models from four architectural families on 300 word pairs from established psycholinguistic corpora, using the Double-Pass Success Rate to control for output biases. We find that most models struggle to maintain coordination once they are required to avoid detectable signals, while one frontier model retains near-perfect performance even after such filtering. We further show that models can direct this capability toward deliberate misdirection, and that coordination is consistently weaker across architectures than within them.