从信号到行为:基于证据的安卓恶意软件检测
From Signals to Behaviors: Evidence-Based Android Malware Detection
浏览论文内容
中文总结 AI 辅助
研究安卓恶意软件检测问题,提出Praxis方法,通过假设-确认-判断流程,从静态信号出发结合程序分析和上下文判断恶意行为。实验证明该方法检测性能优,能降低误报率并恢复细粒度行为,各阶段对性能均有贡献。
中文摘要 AI 辅助
安卓恶意软件仍是持续威胁,准确检测是长期未决问题。应用是否恶意取决于实际行为及上下文,非表面信号。现有检测器基于行为代理判断,导致误判。本文提出Praxis,将检测构建为假设-确认-判断流程,从粗粒度静态信号假设候选行为,通过程序分析用代码证据确认,结合用户意识、应用功能上下文等判断。实验表明,Praxis在三种具有挑战性的设置下优于七个基线,总体检测性能最佳,在高权限良性应用上大幅降低误报率,还能恢复细粒度恶意行为。消融研究表明流程各阶段对最终性能都有贡献。
英文摘要
Android malware remains a persistent threat, and detecting it accurately is a long-standing open problem. Whether an app is malicious depends on what it actually does and the context in which it does it, not on the surface signals it happens to exhibit. Existing detectors instead reason about proxies for behavior, such as learned features or local code slices, and flag whatever deviates from these proxies as malicious. But deviation is not maliciousness: benign apps that merely look unusual are over-flagged, evolving malware that looks ordinary slips through. We argue that detection should be behavior-oriented: recover an app's potentially malicious behaviors and judge which are truly malicious. To realize this, we present Praxis, which structures detection as a hypothesize-confirm-judge pipeline: it hypothesizes candidate behaviors from coarse static signals, confirms each by grounding it in code evidence verified with program analysis, and judges the confirmed behaviors in context: the user's awareness, the app's functional context, and how they compose into an attack. For a malicious app, Praxis returns a verdict and the supported behaviors. We evaluate Praxis against seven baselines across three challenging settings. It achieves the best overall detection performance (87.4% F1), outperforming the baselines by 18.6-34.8 percentage points. On high-permission benign apps, it reduces the false-positive rate to 13.0%, a reduction of 41.1-67.0 percentage points compared with the baselines. Beyond binary detection, Praxis recovers fine-grained malicious behaviors at 87.3% F1, outperforming prior behavior-level approaches by 56.5-73.4 percentage points. Ablation studies show that each stage of the pipeline contributes to the final performance.