arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.12228cs.SE

基于启发式方法与开源大语言模型(LLMs)从源代码中自动提取领域模型的研究

Towards Automated Domain Model Extraction from Source Code using Heuristics and Open-Source LLMs

Alessandra Mancas, Mounir Ammam, Hyacinth Ali, Kevin Delcourt, Houari Sahraoui

首次发表
浏览论文内容

中文总结 AI 辅助

本文提出结合启发式方法与开源LLMs的自动领域模型提取方法,克服上下文限制,在10个项目数据集上获高F1分数,适用于隐私敏感的工业逆向工程场景。

中文摘要 AI 辅助

大语言模型(LLMs)近期展现出强大的代码理解能力,使其有望从源代码中逆向工程领域模型。然而,最先进的专有LLMs因隐私和保密约束无法在许多工业场景中使用,而可本地运行的紧凑型开源LLMs受限于上下文窗口,无法直接处理大型代码库。本文提出一种使用轻量、可本地部署LLMs从源代码自动提取领域模型的方法,该方法结合结构与语义启发式规则及基于LLM的迭代推理,以克服上下文限制。通过逐步分析排名后的代码元素子集,该方法识别领域概念并细化领域边界,无需完整系统上下文。在包含10个项目的数据集上,每个项目均包含经整理的领域模型及其对应实现,本方法取得了较高的F1分数,同时可完全在可本地部署的LLMs上执行,这使其特别适用于隐私敏感型工业环境中的逆向工程任务。

英文摘要

Large language models (LLMs) have recently shown strong capabilities for code understanding, making them promising for reverse engineering domain models from source code. However, state-ofthe- art proprietary LLMs cannot be used in many industrial contexts due to privacy and confidentiality constraints, while compact open-source LLMs that can run locally are limited by their context window and cannot process large code bases directly. In this paper, we propose an automated approach to extract domain models from source code using lightweight, locally deployable LLMs. Our method combines structural and semantic heuristics with iterative LLM-based reasoning to overcome context limitations. By progressively analyzing ranked subsets of code elements, the approach identifies domain concepts and refines domain boundaries without requiring full-system context. Our approach achieves high F1-scores on a dataset of ten projects, each comprising a curated domain model and its corresponding implementation, while remaining fully executable on locally deployable LLMs. This makes it particularly suitable for reverse engineering tasks in privacy-sensitive industrial environments.

补充信息

↑