无文档移动数据库中的结构推断:用于评估数字取证中智能体推理能力的可复现基准
Structural Inference in Undocumented Mobile Databases: A Reproducible Benchmark for Evaluating Agentic Reasoning in Digital Forensics
浏览论文内容
中文总结 AI 辅助
本研究构建了可复现基准,评估智能体对无文档移动数据库的结构推断能力,发现其在规则模式下可靠,模式模糊时性能骤降,需额外验证以确保推断关系可作为可靠取证证据。
中文摘要 AI 辅助
智能体大语言模型正越来越多地用于数字取证分析,但它们推断无文档移动应用数据库内部关系结构的能力仍未被充分理解。在取证场景中,结构不正确的推断会产生看似合理但证据不成立的结果。本研究将智能体结构推断作为一种独立能力进行评估,将执行成功和结构正确性作为两个不同的评估维度。研究考察智能体在仅获得原始数据库和自然语言调查提示时,如何重建表关系、链接属性以及可执行的连接路径。我们将固定的确定性评估流程应用于两个对比鲜明的SQLite仓库:具有稳定标识符传播的Android短信数据库,以及具有不规则模式、临时标识符和多态关系的Snapchat数据库。利用专家验证的SQL基准真值,我们评估:(i)推断关系链接的结构正确性;(ii)多表推理下的执行一致性;(iii)当执行成功但关系解释与专家基准真值存在偏差时,推断结构的鲁棒性和失败模式。评估独立于语义解释进行,完整查询和执行轨迹见附录。结果表明,结构推断在规则模式中保持可靠,但随着模式模糊性增加会急剧下降,频繁产生结构上合理但执行成功的错误连接。这些发现阐明了模式无关的智能体推理可在哪些方面支持取证分析,其鲁棒性在现实模式不规则性下如何下降,以及为何在将推断关系视为可靠证据前仍需额外验证。
英文摘要
Agentic large language models are increasingly used in digital forensic analysis, yet their ability to infer relational structure inside undocumented mobile application databases remains poorly understood. In forensic contexts, structurally incorrect inferences can yield results that appear plausible while remaining evidentially unsound. This work evaluates agentic structural inference as an isolated capability, treating execution success and structural correctness as distinct evaluation axes. It examines how an agent reconstructs table relationships, linking attributes, and executable join paths when given only a raw database and a natural-language investigative prompt. We apply a fixed, deterministic evaluation pipeline to two contrasting SQLite repositories: Android's SMS database with stable identifier propagation, and Snapchat's database with irregular schemas, ephemeral identifiers, and polymorphic relationships. Using expert-verified SQL ground truth, we evaluate (i) structural correctness of inferred relational links, (ii) execution coherence under multi-table reasoning, and (iii) robustness and failure modes of inferred structure when execution succeeds but relational interpretation diverges from expert ground truth. Evaluation is performed independently of semantic interpretation, with full queries and execution traces provided in the Appendix. Results show that structural inference remains reliable in regular schemas but degrades sharply as schema ambiguity increases, frequently producing structurally plausible yet incorrect joins that execute successfully. These findings clarify where schema-agnostic agentic reasoning can support forensic analysis, how its robustness degrades under realistic schema irregularities, and why additional verification remains essential before inferred relationships can be treated as reliable evidence.