发表机构
Rochester Institute of Technology; Gonzaga University(罗切斯特理工学院; 贡萨加大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
该研究针对LLM在网络安全应用中证据归因追踪的缺口,提出拓扑归因距离(TAD)方法,可自适应识别对LLM输出关键的事件日志,为证据验证提供可解释追踪。
AI 中文摘要
大型语言模型(LLM)正越来越多地部署在网络安全操作中,以协助网络安全分析师针对新兴威胁快速做出决策。然而,在网络安全领域使用LLM时必须满足一个主要标准,即对生成输出的信任。随着智能体AI被集成到运营系统中,强大的证据归因和来源追踪技术对于追踪模型生成内容的起源至关重要。当自主智能体做出决策(无论正确与否)时,能够回溯决策链的能力至关重要,因为没有这种能力,团队就无法确定数据的哪一部分导致了模型生成。现有方法往往难以区分复杂且高度相似的证据源,例如网络事件日志。这揭示了一个关键缺口:当前方法无法充分捕捉检索到的证据与生成响应之间的整体几何关系,以实现可靠的证据验证。为了弥合这一缺口,我们受拓扑学启发,提出了拓扑归因距离(Topological Attribution Distance,TAD),以表征并捕获输出的全局几何形状及其相对于检索日志的变化。换句话说,如果特定源日志的嵌入在嵌入空间中显著改变了模型响应的几何形状,则表明该日志是模型生成响应的关键源。因此,TAD由片段级消融归因提供支持,用于调查实际网络攻击的事件日志。我们展示了TAD如何以自适应方式找到对LLM输出贡献最大的日志,这可以基于每个LLM的隐藏状态提供可解释且可信的追踪,以理解检索到的日志在几何上如何不同地影响模型生成,并在网络安全和智能体AI工作流中提供证据验证。
英文摘要
Large Language Models (LLMs) are increasingly being deployed in cybersecurity operations to assist cybersecurity analysts with rapid decision-making against emerging threats. However, there is a main criteria that must be met when using LLMs in cybersecurity, that is, trust in the generated outputs. As Agentic AI is integrated into operational systems, a robust evidence attribution and provenance tracking technique is essential to trace the origins of model generations. When autonomous agents make a decision (right or wrong), the ability to trace back through the decision chain is critical, as without it, teams cannot identify which segment of the data caused the model generation. Existing methods often struggle to distinguish among complex and highly similar evidence sources, such as cyber incident logs. This reveals a key gap: current approaches do not adequately capture the holistic geometric relationship between the retrieved evidence and the generated response for reliable evidence verification. To bridge this gap, we propose Topological Attribution Distance (TAD), inspired by Topology, to characterize and capture the global geometric shape of an output and its changes against its retrieved logs. In other words, if the embeddings of a specific source log drastically changes the geometry of the model's response in the embedding space, this suggests that such log is a critical source for the model's generated response. Therefore, TAD is powered by segment-level ablation attribution to investigate incident logs of an actual cyberattack. We demonstrate how TAD finds the most attributed logs on LLM outputs in an adaptive manner. This can provide an explainable and trustworthy tracing based on each LLM's hidden state to understand how geometrically different retrieved logs influence the model generation, and provide evidence verification in cybersecurity and Agentic-AI workflows.