面向大语言模型增强的Android污点分析
Towards LLM-Enhanced Android Taint Analysis
AI总结:
本文探究商用LLM能否有效推理Android应用的污点流,提出智能体交互策略,在DroidBench基准上Gemini-3 Flash的F1分数达0.96(优于FlowDroid的0.55),还能识别真实应用中FlowDroid未发现的潜在数据泄露,可补充传统静态污点分析。
AI中文摘要:
污点分析是检测Android应用中敏感数据泄露的基础技术。然而,传统静态工具(如FlowDroid)因准确建模Android框架的复杂性仍面临公认挑战。本文研究商用大语言模型(LLM)能否有效推理Android应用中的污点流。我们的初步方法采用智能体交互策略,使LLM能迭代探索代码并推理数据流。我们在DroidBench基准上针对FlowDroid开展初步评估,结果显示我们的方法优于基线:Gemini-3 Flash的F1分数达0.96,而FlowDroid为0.55。尤其在组件间通信(0.95 vs. 0.17)、隐式流(0.94 vs. 0.00)和反射(1.00 vs. 0.50)等FlowDroid通常表现不佳的挑战性类别中,我们的方法取得了提升。在一小部分真实应用上,基于LLM的方法还识别出FlowDroid未报告的额外潜在数据泄露。这些初步发现表明,LLM推理可有效补充传统静态污点分析,为未来研究混合LLM增强的污点分析管道提供了动机。
英文摘要:
Taint analysis is a fundamental technique for detecting sensitive data leaks in Android apps. However, traditional static tools, such as FlowDroid, still face well-known challenges due to the complexity of accurately modeling the Android framework. In this paper, we investigate whether off-the-shelf Large Language Models (LLMs) can effectively reason about taint flows in Android apps. Our preliminary approach relies on an agentic interaction strategy, enabling the LLM to iteratively explore code and reason about data flows. We conduct an initial evaluation on the DroidBench benchmark against FlowDroid, where our approach outperforms the baseline: Gemini-3 Flash achieves an F1-score of 0.96, compared to 0.55 for FlowDroid. In particular, we observe improvements in challenging categories such as inter-component communication (0.95 vs. 0.17), implicit flows (0.94 vs. 0.00), and reflection (1.00 vs. 0.50), where FlowDroid typically struggles. On a small set of real-world apps, the LLM-based approach also identifies additional potential data leaks not reported by FlowDroid. These preliminary findings suggest that LLM reasoning may effectively complement traditional static taint analysis, motivating future research on hybrid LLM-enhanced taint analysis pipelines.