发表机构
Cranberry-Lemon University; University of the Witwatersrand(克莱恩伯里-柠檬大学; 威特沃特斯兰德大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对现有LLM护栏难以识别隐性成长风险的问题,提出SaplingGuard多智能体护栏,通过用户画像、上下文风险评估和提示优化,在SaplingBench上将Major Hit率提升至63.7%,有害响应率降至5.27%,有效保障青少年交互安全。
AI 中文摘要
随着青少年在日常生活中越来越多地使用大语言模型(LLM),确保其响应安全且符合发展适宜性变得至关重要。然而,现有的LLM护栏主要针对孤立提示或响应中的显性有害内容,在识别隐性的、依赖上下文的成长风险方面效果不佳。为解决这一局限,我们提出了SaplingGuard,一种即插即用、画像感知和对话感知的护栏,无需修改下游模型参数。SaplingGuard将青少年安全干预分解为三个专门智能体,分别负责用户画像构建、上下文感知风险评估和意图保持的提示优化。这些智能体共同利用当前提示、先前对话和结构化用户特征来识别上下文风险并引导下游响应生成。我们在SaplingBench上评估了SaplingGuard,该基准包含276个三轮对话,涵盖七类成长风险。在十种青少年画像条件和九种开源及闭源下游LLM中,画像感知检索将Major Hit率从50.8%提升至63.7±1.1%。端到端干预进一步将平均有害响应率从17.10%降低至5.27%,并将平均安全得分从0.7017提升至1.0043。这些结果表明,用户画像和对话上下文为识别隐性成长风险提供了互补信号,且SaplingGuard可作为青少年与大语言模型交互的有效外部安全层。
英文摘要
As adolescents increasingly use LLMs in everyday life, ensuring safe and developmentally appropriate responses has become essential. However, existing LLM guardrails primarily target explicit harmful content in isolated prompts or responses and are less effective at identifying implicit, context-dependent developmental risks. To address this limitation, we propose SaplingGuard, a plug-and-play, profile-aware and dialogue-aware guardrail that requires no modification to downstream model parameters. SaplingGuard decomposes adolescent safety intervention into three specialized agents for user profile construction, context-aware risk assessment, and intent-preserving prompt optimization. Together, these agents leverage the current prompt, preceding dialogue, and structured user characteristics to identify contextual risks and guide downstream response generation. We evaluate SaplingGuard on SaplingBench, which contains 276 three-turn dialogues spanning seven categories of developmental risk. Across ten adolescent profile conditions and nine open- and closed-source downstream LLMs, profile-aware retrieval improves the Major Hit rate from 50.8% to 63.7+/-1.1%. End-to-end intervention further reduces the average harmful response rate from 17.10% to 5.27% and increases the average safety score from 0.7017 to 1.0043. These results show that user-profile and dialogue context provide complementary signals for identifying implicit developmental risks, and that SaplingGuard can serve as an effective external safety layer for adolescent-LLM interaction.