用于安全开放式探索的异构智能体群组及运行时约束记忆
Heterogeneous Agent Cohorts for Safe Open-Ended Exploration with Runtime Constraint Memory
浏览论文内容
中文总结 AI 辅助
研究针对语言模型智能体安全与探索困境,提出将关注点分离到专门角色,通过蒙特卡洛树搜索编译失败为‘伤疤’约束补丁,在空间语义沙盒实验中,该方法能到达远程目标、防止违规、减少令牌消耗,还能在资源约束下降低成本。
中文摘要 AI 辅助
如今的语言模型智能体处于尴尬境地。用静态安全指令限制它们,它们很少能超越明显的范围;给予它们使用工具和多智能体辩论的自由,安全违规很快就会出现。我们将关注点分离到专门的角色中,破坏者生成非常规提议,验证者在工具网关执行严格的运行时检查,中介引入遥远但相关的类比。失败不会被丢弃,而是通过蒙特卡洛树搜索编译成我们称为‘伤疤’的紧凑、带符号的约束补丁。这些补丁在本地缓存并被未来群组继承,将重复失败转化为可复用、低成本的运行时约束。在空间语义沙盒中(N = 20次运行,p < 0.01),我们的群组能到达辩论失败的远程目标,验证者防止所有执行的违规行为,‘伤疤’通过避免冗余验证检查将令牌消耗减少15.1%。此外,基于信用的通信分配分数限制出站带宽,在资源约束下将总体令牌成本降低55.9%。
英文摘要
LLM agents today are caught in an awkward bind. Lock them down with static safety instructions and they rarely venture beyond the obvious; give them free reign with tools and multi-agent debate, and safety violations quickly follow. Rather than forcing a single model to juggle both creativity and caution, we separate the concerns across specialized roles. A Disrupter generates unconventional proposals, a Validator enforces hard runtime checks at the tool gateway, and a Broker pulls in distant but relevant analogies. Failures are not discarded -- they are compiled, via MCTS, into compact, signed constraint patches we call Scars. These patches are cached locally and inherited by future cohorts, turning repeated failures into reusable, low-cost runtime constraints. In a spatial-semantic sandbox (N=20 runs, p<0.01), our cohort reaches remote targets where debate fails, the Validator prevents all executed breaches, and Scars reduce token consumption by 15.1% by avoiding redundant validator checks. Furthermore, credit-based Communication Allocation Scores (CAS) restrict outbound bandwidth, reducing overall token costs by 55.9% under resource constraints.
发表机构
- Google DeepMind(谷歌DeepMind)
机构由 AI 辅助整理,请以论文原文为准。