通过提示优化在指称博弈中实现冻结大语言模型间的通信
Communication between Frozen Large Language Models via Prompt Optimization in a Referential Game
中文总结 AI 辅助
本研究通过提示优化使冻结大语言模型在指称博弈中建立通信,采用独立优化器重写提示,实现共享代码泛化,并验证了位值通信协议的可审计性。
中文摘要 AI 辅助
我们研究了来自不同提供商、具有不同分词器、通过API端点访问的两个冻结大语言模型之间的通信。这两个模型进行一个指称博弈:一个模型看到一个对象,并用小字母表上的短固定长度消息对其进行描述;另一个模型必须从候选集中选出该对象。两个模型的权重均不更新。每个智能体的提示由独立的提示优化器重写,该优化器的反思模型读取该智能体的带评分交互。在位置设置中,优化后的提示携带一种共享代码,该代码在保留的测试对象上泛化,优于测量的无代码本基线,包括在移除记忆窗口时。在第二种设置中,独立的逐字母块不再适合消息,但整体对象的位值代码可以。基础系统无法建立可靠的通信:发送方难以保持单射规则,接收方在视野中确认的示例太少。发送方碰撞惩罚、成功交互的保留以及顺序优化在某些运行中实现了成功的位值通信。结果因运行和反思模型而异。在成功运行中,协议被写入优化后的提示中,可以直接读取和审计。
英文摘要
We study communication between two frozen large language models from different providers, with different tokenizers, accessed through their API endpoints. The two play a referential game: one sees an object and describes it in a short fixed-length message over a small alphabet; the other must pick that object out of a candidate set. Neither model's weights are updated. Each agent's prompt is rewritten by an isolated prompt optimizer whose reflection model reads that agent's scored interactions. In the positional setting, optimized prompts carry a shared code that generalizes to held-out objects above a measured no-codebook baseline, including when the memory window is removed. In a second setting, independent per-letter blocks no longer fit within the message, although a whole-object place value code does. The base system fails to establish reliable communication: the sender struggles to retain an injective rule, and the receiver has too few confirmed examples in view. A sender collision penalty, retention of successful interactions, and sequential optimization enable successful place value communication in some runs. Outcomes vary across runs and reflection models. In successful runs, the protocol is written into the optimized prompts, where it can be read and audited directly.