发表机构
University of Florence; Nanyang Technological University(佛罗伦萨大学; 南洋理工大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出个性化轨迹级安全框架,通过多模态轨迹采样和生成流网络,在每轮对话中筛选并选择保留最多安全延续的策略,以降低拟人化AI交互中的累积风险,同时保持有用性。
AI 中文摘要
拟人化人工智能系统日益记住个人细节、展现同理心,并被当作社会伙伴来互动,从而产生了随着用户与系统关系随时间演变而出现的新型风险。现有的安全措施主要在单个对话轮次层面运作,无法判断一系列看似可接受的互动是否在累积性地将特定用户推向伤害。本文引入了个性化轨迹级安全(personalized trajectory-level safety),该框架将关系安全视为一个序贯决策问题,其状态是一个潜在升级状态(latent escalation state),该状态从用户的消息中推断,并受系统响应的影响。在每一轮中,筛选步骤首先丢弃任何在用户的所有合理模型下都无法保留至少一条安全互动延续的响应策略。在剩余策略中,我们将动作选择表述为多模态轨迹采样,并使用生成流网络(Generative Flow Network)生成多样化的未来演变,其生成比例与它们的合理性、安全性和效用成正比。然后,系统选择保留最大比例的安全且有用的延续的策略。我们在模拟中评估该框架,模拟基于真实人机对话互动的统计数据校准,并使用来自公开基准的响应策略。结果表明,轨迹感知决策大幅降低了有害状态的频率,同时保持了有用的互动。这项工作将拟人化AI的安全从响应级过滤重新定义为对人类-AI关系未来演变的个性化控制。源代码可在以下网址获取:此 https URL。
英文摘要
Anthropomorphic artificial intelligence systems increasingly remember personal details, display empathy, and are engaged with as social counterparts, creating forms of risk that emerge from the evolution of the user-system relationship over time. Existing safeguards largely operate at the level of individual conversational turns and cannot determine whether a sequence of seemingly acceptable interactions is cumulatively moving a particular user toward harm. This paper introduces personalized trajectory-level safety, a framework that treats relational safety as a sequential decision problem over a latent escalation state inferred from the user's messages and influenced by the system's responses. At each turn, a screening step first discards any response strategy that does not preserve at least one safe continuation of the interaction under every plausible model of the user. Among the remaining strategies, we formulate action selection as multimodal trajectory sampling, and use a Generative Flow Network to generate diverse future evolutions in proportion to their plausibility, safety, and utility. The system then selects the strategy that preserves the largest fraction of safe and useful continuations. We evaluate the framework in simulation, calibrated on statistics reported for real human-chatbot interactions, and using response strategies derived from public benchmarks. Results show that trajectory-aware decision making substantially reduces the frequency of harmful states while keeping helpful interaction. This work reframes safety for anthropomorphic AI from response-level filtering to personalized control over the future evolution of human-AI relationships. The source code is available at https://github.com/benedettapicano/ANTHROPOMORPHIC_SAFETY_TRAJ.