arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.21107cs.AIcs.SE

大型语言模型在软件工程与软件安全交叉领域:一项以证据为中心的结构化调查与研究议程

Large Language Models at the Intersection of Software Engineering and Software Security:An Evidence-Centered Structured Survey and Research Agenda

  • Nanjing Liancheng Intelligent Technology Group(南京连城智能科技集团)

机构由 AI 辅助整理,请以论文原文为准。

Wei Lin, Tao Zhou, Zhaofei Xie, Changgui Hong

AI总结:

本研究调查了LLMs在软件工程与软件安全交叉领域的进展,提出保证框架,识别有效性威胁与最低报告协议,制定研究议程,主张以任务适配证据而非单一基准评判模型能力。

AI中文摘要:

大型语言模型(LLMs)正从代码补全向仓库级智能体演进,这类智能体可检索上下文、编辑文件、执行工具并参与对安全敏感的工作流。然而,关于这些系统的证据仍存在分歧:软件工程评估聚焦于功能任务完成,而软件安全评估则聚焦于漏洞检测、安全生成或利用导向的验证。这项以证据为中心的结构化调查综合了截至2026年5月31日在软件工程任务、软件安全任务、适配机制、工件粒度及评估设计方面的代表性工作。除任务分类外,我们引入了一种保证框架,将功能正确性、安全性、操作可靠性、证据来源及智能体权限区分开来。综述显示,执行反馈与仓库访问可大幅提升工程任务完成效果,但本身并不能保障安全性;相反,静态分析标签或漏洞分类评分也极少能确立可部署的正确性。我们识别出反复出现的有效性威胁——弱测试神谕、重复及时间上泄露的数据、变化的智能体测试工具、仅基于代理的安全检查,以及未充分报告的预算与人工干预——并推导了用于跨研究比较的最低报告协议。由此得出的研究议程优先考虑兼具安全性与功能性的基准、仓库级威胁模型、校准后的人工监督、长期可维护性证据,以及可复现的智能体评估。核心结论是,模型能力应根据任务适配的证据所支持的保证案例来评判,而非单一基准分数。

英文摘要:

Large Language Models (LLMs) are moving from code completion toward repository-scale agents that retrieve context, edit files, execute tools, and participate in security-sensitive workflows. The evidence for these systems, however, remains divided between software engineering evaluations centered on functional task completion and software security evaluations centered on vulnerability detection, secure generation, or exploit-oriented validation. This evidence-centered structured survey synthesizes representative work available through May 31, 2026 across software engineering tasks, software security tasks, adaptation mechanisms, artifact granularity, and evaluation design. In addition to a task taxonomy, we introduce an assurance framework that separates functional correctness, security, operational reliability, evidence provenance, and agent authority. The review shows that execution feedback and repository access can substantially improve engineering task completion, but do not by themselves establish security; conversely, static-analysis labels or vulnerability-classification scores rarely establish deployable correctness. We identify recurring validity threats--weak test oracles, duplicated and temporally leaked data, changing agent harnesses, proxy-only security checks, and under-reported budgets and human intervention--and derive a minimum reporting protocol for cross-study comparison. The resulting research agenda prioritizes jointly secure-and-functional benchmarks, repository-scale threat models, calibrated human oversight, longitudinal maintainability evidence, and reproducible agent evaluation. The central conclusion is that model capability should be judged as an assurance case supported by task-appropriate evidence, rather than by a single benchmark score.

↑