Adversarial Stress Testing of Role-Playing Language Agents using Multi-Agent Evaluation
基于多智能体评估的角色扮演语言智能体的对抗性压力测试
专题命中 机器人数据与评测 :manipulation(abstract,abstract_cn);分类 cs.AI
AI总结 本研究提出模块化多智能体平台,通过多轮对话对角色扮演语言智能体开展对抗性压力测试,揭示单策略测试无法发现的故障模式,相关成果作为开源平台发布以支持AI安全。
Comments 8 pages, 1 figure, 7 tables; accepted and presented at ADScAI Conference 2026, University of Moratuwa, Sri Lanka