在安全调查中使用和评估大语言模型(LLM)作为代理专家的框架:可靠性、偏差及启示
A Framework for Using and Evaluating LLMs as Surrogate Experts in Security Surveys: Reliability, Bias, and Implications
浏览论文内容
中文总结 AI 辅助
本研究提出了评估LLM作为安全调查代理专家的框架,发现LLM虽内部一致但与专家响应存在系统性偏差,可用于试点和假设生成但不能替代专家征询。
中文摘要 AI 辅助
专家调查广泛应用于安全研究中,用于研究从业者的工作流程和决策,但招募领域专家——尤其是安全运营中心(SOC)的分析师,他们面临高工作量、职业倦怠和保密性限制——十分困难,且往往导致样本量较小。大语言模型(LLM)提供了一种有吸引力的替代方案,可大规模生成合成响应,但关于这类代理参与者何时可靠的指导意见甚少。我们提出了一种用于评估LLM作为专家调查受访者替代品或补充的方法学框架。利用SOC专业人员的响应,我们在多个模型和提示设置下比较了基于角色设定(persona-based)和聚合式(aggregate)的LLM生成答案。我们测量了稳定性、模型间一致性以及与人类响应的对齐程度。结果表明,尽管LLM生成的答案内部一致,但它们与专家答案存在系统性偏差,表现出方差降低、集中趋势偏差和意见同质化。本研究为安全研究界提供了关于LLM生成调查响应的恰当使用方式和局限性的方法学证据与实践指导。我们得出结论,LLM对于试点研究和假设生成是有用的,但不能替代专家征询,并讨论了使用LLM增强型调查的研究者所面临的启示。
英文摘要
Expert surveys are widely used in security research to study practitioner workows and decision-making, yet recruiting domain experts - especially in Security Operations Centres (SOCs), where analysts face high workload, burnout and confidentiality constraints - is difficult and often results in small samples. Large language models (LLMs) oer an appealing alternative by generating synthetic responses at scale, but little guidance exists on when such surrogate participants are reliable. We present a methodological framework for evaluating LLMs as substitutes or supplements to expert survey respondents. Using responses from SOC professionals, we compare persona-based and aggregate LLM-generated answers across multiple models and prompting settings. We measure stability, inter-model agreement and alignment with human responses. Our results show that although LLMs produce internally consistent answers, they systematically diverge from experts, exhibiting reduced variance, central tendency bias and homogenised opinions. This work contributes methodological evidence and practical guidance to the security research community on the appropriate use and limitations of LLM-generated survey responses. We conclude that LLMs are useful for piloting and hypothesis generation but not for replacing expert elicitation, and we discuss implications for researchers using LLM-augmented surveys.
发表机构
- SnT, University of Luxembourg(卢森堡大学SnT研究院)
- Cyber-Defence Campus, armasuisse Science and Technology (S+T)(armasuisse科技公司网络防御园区)
机构由 AI 辅助整理,请以论文原文为准。