Heterogeneous Agent Cohorts for Safe Open-Ended Exploration with Runtime Constraint Memory
用于安全开放式探索的异构智能体群组及运行时约束记忆
机构 * Google DeepMind(谷歌DeepMind)
AI总结 研究针对语言模型智能体安全与探索困境,提出将关注点分离到专门角色,通过蒙特卡洛树搜索编译失败为‘伤疤’约束补丁,在空间语义沙盒实验中,该方法能到达远程目标、防止违规、减少令牌消耗,还能在资源约束下降低成本。
Comments 12 pages, 1 figure, 12 tables