发表机构
Guangzhou Health Science College(广州卫生职业技术学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对模块化LLM安全智能体的轨迹风险验证问题,提出生成树替代方案,在两阶段入侵检测实验中实现92.7%±2.4%的平均轨迹覆盖率,量化了联合访问缺失的成本。
AI 中文摘要
自主安全智能体以分段流水线形式运行,例如先对网络流量进行分类,再将攻击归因于特定技术。拆分式共形预测为每个阶段提供有限样本覆盖率,但部署需要对完整链路提供轨迹级别的保证。当各阶段独立训练和校准时,这些保证不会自动组合。Bonferroni分配是无分布的,但在误差相关时较为保守。本文表明,将其自然扩展到三个或更多阶段的成对相关方法是无效的,因为它给出的是下界而非上界,因此推导了一种有效的生成树替代方案。本文区分了阶段是否相关与审计样本是否足够大以验证这种相关性,并给出了匹配的上界和信息论下界样本复杂度界。本文还表明,由粗到细的标签选择可在无需学习相关性的情况下创建近乎完美的测量相关性。在包含6个开源LLM和2个数据集的两阶段入侵检测流水线中,消除该人工产物可将测量相关性从接近1降至0-0.78。一旦审计达到所需样本量,对轨迹故障的直接审计比Bonferroni方法紧13.7%,但样本量不足时效果更差。使用每阶段证书和成对重叠界的模块化证书产生了0.6%的正平均增益,量化了缺乏联合访问的成本。同模型、跨模型和置换配对测试表明,残余相关性反映了共享样本难度,而非共享模型表示。在α=0.10下,12种配置的平均轨迹覆盖率为92.7%±2.4%。在跨数据集部署下,即使准确率保持78%,单步未覆盖率也达到100%,表明分布偏移在原始准确率之前破坏了校准的置信度。
英文摘要
Autonomous security agents operate as staged pipelines, such as classifying network traffic and then attributing attacks to a specific technique. Split conformal prediction gives each stage finite-sample coverage, but deployment requires a trajectory-level guarantee across the full chain. These guarantees do not compose automatically when stages are independently trained and calibrated. Bonferroni allocation is distribution-free but conservative under correlated errors. We show that a natural pairwise-correlation extension to three or more stages is invalid because it gives a lower rather than an upper bound, and derive a valid spanning-tree alternative. We distinguish whether stages are dependent from whether an audit sample is large enough to certify that dependence, and give matching upper and information-theoretic lower sample-complexity bounds. We also show that coarse-to-fine label selection can create near-perfect measured correlation without learned dependence. On a two-stage intrusion-detection pipeline across 6 open LLMs and 2 datasets, removing this artifact reduces measured correlation from near 1 to 0-0.78. A direct audit of trajectory failure becomes 13.7% tighter than Bonferroni once the audit reaches the required sample size, but is worse when undersized. A modular certificate using per-stage certificates and a pairwise overlap bound yields a positive average gain of 0.6%, quantifying the cost of lacking joint access. Same-model, cross-model, and permuted-pairing tests show that residual dependence reflects shared sample difficulty, not shared model representations. Average trajectory coverage across 12 configurations is 92.7% +/- 2.4% at alpha = 0.10. Under cross-dataset deployment, single-step miscoverage reaches 100% even when accuracy remains 78%, showing that distribution shift destroys calibrated confidence before raw accuracy.
Comments18 pages, 11 tables; preprint