AI 中文总结
研究大语言模型翻译对手模拟程序的评估问题,核心方法是基于图的结构评估(GBSE),将程序建模为有向属性图并计算归一化图编辑距离,主要贡献是应用于跨平台计划评估,揭示技术差异,框架含多种工具和规则且结果可重现验证。
AI 中文摘要
对手模拟计划描述了使用MITRE ATT&CK技术、权限要求和可观测遥测的多步骤攻击者程序。跨操作系统翻译它们有助于跨平台防御者评估,大语言模型可自动化此任务。但翻译可能仅重命名工具而保留源平台逻辑,二进制评分会高估保真度。基于图的结构评估(GBSE)将每个程序建模为有向属性图,并计算四层的归一化图编辑距离。将GBSE应用于一个29步的ALPHV/BlackCat从Windows到Linux的计划,比较重构的Windows控制与未修改的大语言模型生成的Linux版本。技术和策略结构完全保留,遥测相似度降至0.897,Sigma日志源相似度为1.000。每个状态被分类为中等保真度,最佳综合评分为0.674,未达到0.80的部署阈值。该框架包括二分图编辑距离、遥测意图解析器和49个经过验证的Sigma规则,还揭示了技术层面的差异。结果通过记录输出进行了重现和验证。
英文摘要
Adversary emulation plans describe multi-step attacker procedures using MITRE ATT&CK techniques, privilege requirements, and observable telemetry. Translating them across operating systems supports cross-platform defender evaluation, and large language models (LLMs) can automate this task. However, a translation may only rename tools while retaining source-platform logic, giving defenders little target-platform coverage. Binary scoring can overestimate fidelity because it measures countable features rather than structural, observable, and rule-level equivalence. Graph-Based Structural Evaluation (GBSE) models each procedure as a directed attributed graph and calculates normalized Graph Edit Distance (GED) across four layers: technique, tactic, telemetry class, and Sigma logsource. GBSE was applied to a 29-step ALPHV/BlackCat Windows-to-Linux plan, comparing a reconstructed Windows control with the unmodified LLM-generated Linux version. Technique and tactic structure were fully preserved (GED=0, similarity=1.000). Telemetry similarity fell to 0.897 (GED=3) because three steps contained unmapped or drifting observables, while Sigma logsource similarity was 1.000. Every state was classified as Medium Fidelity, with a best composite score of 0.674. The 0.80 deployment threshold was not reached because technical realism scored 0.43 against the required 0.990. The framework includes bipartite GED, a telemetry-intent parser that converts free text into observable classes, and 49 validated Sigma rules: 19 for Linux and 30 for Windows. The rules provide complete ATT&CK technique coverage and pass validation with zero findings. Additional analysis reveals technique-level divergence, including RDP-based external access mapped to unencrypted exfiltration and credential-store access mapped to remote-system discovery. Results were reproduced and verified against recorded outputs.
CommentsThis technical contribution supports the MITRE white paper titled: Evaluating LLMs for Impact-Faithful Translation of Adversary Behavior Across Operating Systems