感知成本的分层多智能体勒索软件检测与家族归因
Cost-Aware Hierarchical Multi-Agent Ransomware Detection and Family Attribution under Analysis Budgets
浏览论文内容
中文总结 AI 辅助
提出感知成本的分层多智能体系统(HMAS),通过自适应选择模态平衡勒索软件检测与家族归因的准确率和计算成本,在相关任务中取得良好效果且降低了成本与延迟。
中文摘要 AI 辅助
勒索软件检测与家族归因需要分析不同模态,因为它会使用打包、混淆、进程操纵和运行时规避技术。然而,传统多模态方法通常对每个样本使用所有可用模态,导致不必要的计算成本和延迟增加。本文提出一种感知成本的分层多智能体系统(HMAS)用于自适应勒索软件检测。所提架构将专业智能体组织为由元协调器协调的分层域控制器。静态分析作为初始低成本模态,当置信度不足或专业智能体意见不一致时,选择性使用额外的动态和内存模态。成本模型整合模态使用和处理开销,使协调策略能在分析性能与计算成本间取得平衡。本地部署的大语言模型(LLM)为选定的困难案例提供验证,不取代确定性流程。实验评估在二进制勒索软件检测和多类家族归因任务中,将自适应HMAS与仅静态、静态加动态及穷尽分析策略对比。完整HMAS在二进制检测中达到96.57%准确率、0.96 F1值和0.99 ROC-AUC,在家族归因中达到0.90宏F1值。同时,HMAS相比穷尽分析降低43.97%的平均分析成本,除使用LLM的情况外,平均分析延迟也大幅降低。路由分析显示,56.05%的案例仅通过静态证据即可解决,仅4.33%需要完整证据流程。这些发现表明,自适应HMAS可在勒索软件分析中实现准确率与成本的权衡,同时保留对异构和不完整模态的支持。
英文摘要
Sandbox execution and memory forensics are among the most constrained resources in malware triage. Static analysis can scale to millions of files, whereas dynamic and memory analysis require minutes of analyst controlled infrastructure for each sample. Despite this difference, multimodal ransomware detectors often apply every modality to every sample, causing analysis cost and time to verdict to increase linearly with sample volume even when static evidence is already sufficient for a decision. We present a cost aware Hierarchical Multi-Agent System that formulates evidence acquisition as a budgeted sequential decision problem. Specialist agents generate schema validated risk signals for each modality, domain controllers aggregate these signals, and a Meta-Orchestrator begins with static evidence and escalates to dynamic and memory evidence only when confidence is insufficient or agents within a controller disagree. An optional, bounded, locally hosted large language model reviewer can adjust a verdict by at most one tier but cannot replace the deterministic pipeline. Each decision is recorded with a complete provenance trace. In multiple runs over 12439 samples from 16 ransomware families and benign samples, the deterministic HMAS achieves F1 0.93 and macro F1 0.97, resolving 57.95% of cases using static evidence alone, 35.83% after adding dynamic evidence, and only 6.21% through the full pipeline. The average internal analysis cost is 6.65 units, compared with 12 for exhaustive analysis, representing a 44.6% reduction. Standalone leave-one-family-out testing further shows that accuracy on families held out during tuning falls to 0.26 to 0.64 outside the Benign and high support classes. We report these results alongside a cost sensitivity analysis, a partial leave one component out ablation, and a full scale comparison with learned early and late fusion and cascade baselines.