arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

评估对话式网络安全助手中个性化策略的影响

Evaluating the Impact of Personalization in Conversational Cybersecurity Assistants

Lea Duesterwald, Anika Jain, Shreya Kochar, Norman Sadeh

arXiv 2609.17839首次发表:更新:

发表机构

Cornell University; Carnegie Mellon University(康奈尔大学; 卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究评估了四种个性化策略对LLM网络安全助手答案效果的影响,发现基于对话历史的个性化最受青睐,且LLM评估与人工评估趋势一致,为高效比较个性化策略提供了可扩展方法。

AI 中文摘要

用户越来越多地转向大型语言模型来回答各种问题,包括网络安全问题。我们研究了个性化策略如何帮助提高基于LLM的网络安全助手所回答问题答案的有效性。除了准确性之外,我们重点关注答案的可理解性、可操作性,以及最重要的是激励作用,因为用户经常未能遵循网络安全建议。具体而言,我们研究了四种个性化策略,从静态用户画像到基于交互历史的个性化,使用了包含1,045个真实世界网络安全问题的语料库,并进行了一项为期7天的部署,涉及57名参与者和1,066个用户问题。在大规模的基于LLM的自动化评估和人工评估中,基于对话的个性化在感知帮助性和遵循安全建议可能性的比较评分中始终受到青睐。重要的是,在基于LLM的评估中观察到的相对趋势与人工评估获得的趋势一致,这表明基于LLM的评估可以在成本高昂的用户研究之前提供一种可扩展的机制来比较个性化策略。这些结果表明,行为驱动的个性化是LLM驱动的网络安全助手的一个有前景的方向,并强调了在研究个性化语言模型系统时结合基于LLM的评估和人工评估的价值。

英文摘要

Users increasingly turn to Large Language Models to answer a variety of questions, including cybersecurity questions. We study how personalization strategies can help improve the effectiveness of answers to questions asked to an LLM-based cybersecurity assistant. Beyond accuracy, we focus on the understandability, actionability and, most importantly, motivating power of answers, given how often users fail to follow cybersecurity recommendations. Specifically, we investigate four personalization strategies, ranging from static user profiles to interaction-history-based personalization, using a corpus of 1,045 real-world cybersecurity questions and a 7-day deployment involving 57 participants and 1,066 user questions. Across both a large-scale automated LLM-based evaluation and human evaluation, conversation-based personalization is consistently favored in comparative ratings of perceived helpfulness and likelihood of following security advice. Importantly, the relative trends observed in the LLM-based evaluation align with those obtained from human evaluation, suggesting that LLM-based evaluation can provide a scalable mechanism for comparing personalization strategies before costly user studies. These results indicate that behavior-driven personalization is a promising direction for LLM-powered cybersecurity assistants and highlight the value of combining LLM-based and human evaluation when studying personalized language-model systems.

CommentsAccepted as a poster at the HAIPS Workshop at COLM 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑