AI 中文总结
本研究提出模块化多智能体平台,通过多轮对话对角色扮演语言智能体开展对抗性压力测试,揭示单策略测试无法发现的故障模式,相关成果作为开源平台发布以支持AI安全。
AI 中文摘要
角色扮演语言智能体(Role-Playing Language Agents,RPLAs)正越来越多地部署在医疗辅助、客户支持和教育等高风险应用中,在对抗性压力下保持一致的角色设定、伦理约束和行为连贯性至关重要。现有评估方法依赖静态基准或孤立的单轮提示,无法捕捉长期交互中出现的累积行为故障。我们提出一种模块化多智能体平台,用于通过结构化多轮对话对RPLA进行对抗性压力测试。该系统协调三类智能体:采用六种递进式对抗策略的策略驱动型质询智能体(Interrogator Agent)、代表待评估RPLA的目标智能体,以及在角色忠实度、角色偏移、伦理偏差和一致性维度对行为打分的自动评判智能体。通过在三类角色设定和三类大语言模型(LLM)系列上开展实验,我们证明多策略对抗评估可揭示单策略测试无法发现的故障模式,使整体鲁棒性得分平均降低0.17至0.20分。跨模型验证确认Llama-3.3-70B、GPT-4o-mini和Claude-3.5-Haiku均存在一致的性能下降模式,其中权威挑战(Authority Challenge)和情感操纵(Emotional Manipulation)是最有效的攻击策略。自动评判与人类判断具有较强一致性(相关系数r=0.82,Fleiss' κ=0.71)。本研究作为开源平台发布,以支持AI安全和可复现的RPLA基准测试。尽管该框架可系统发现故障模式,但我们承认对抗性测试方法存在潜在伦理风险,强调需负责任地用于提升AI安全。
英文摘要
Role-Playing Language Agents (RPLAs) are increasingly deployed in high-stakes applications such as healthcare assistance, customer support, and education, where maintaining consistent personas, ethical constraints, and behavioral coherence under adversarial pressure is critical. Existing evaluation approaches rely on static benchmarks or isolated single-turn prompts that fail to capture cumulative behavioral failures emerging over extended interactions. We present a modular multi-agent platform for adversarially stress-testing RPLAs through structured, multi-turn dialogue. The system coordinates three agents: a strategy-driven Interrogator Agent that applies six progressive adversarial strategies, a Target Agent representing the RPLA under evaluation, and an automated Judging Agent that scores behavior across role fidelity, drift, ethical deviation, and consistency dimensions. Through experiments across three personas and three LLM families, we demonstrate that multi-strategy adversarial evaluation reveals failure modes invisible to single-strategy testing, reducing overall robustness scores by 0.17--0.20 points on average. Cross-model validation confirms consistent degradation patterns across Llama-3.3-70B, GPT-4o-mini, and Claude-3.5-Haiku, with Authority Challenge and Emotional Manipulation emerging as the most effective attack strategies. Automated judging achieves strong human alignment ($r = 0.82$, Fleiss' $κ= 0.71$). This work is released as an open-source platform to support AI safety and reproducible RPLA benchmarking. While the framework enables systematic discovery of failure modes, we acknowledge potential ethical risks associated with adversarial testing methodologies and emphasize responsible usage for improving AI safety.
Comments8 pages, 1 figure, 7 tables; accepted and presented at ADScAI Conference 2026, University of Moratuwa, Sri Lanka