arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.04667cs.SE

正确性、置信度与上下文:在人工智能时代构建软件保障

Correctness, confidence, and context: Framing software assurance in the AI age

Mary Shaw

首次发表
浏览论文内容

中文总结 AI 辅助

探讨软件工程中正确性相关挑战,指出传统方法与生成式人工智能在软件保障上的差异,呼吁系统思考保障技术,做出明智选择。

中文摘要 AI 辅助

软件工程与“正确性”关系复杂。生成式人工智能为保障带来新维度,其基于统计。传统软件工程通过严格推理等建立置信度,而生成式人工智能结果是“可能近似正确”的预测,限制了保障,且无法纳入隐性知识。我们需系统思考保障技术,明智选择方法组合。

英文摘要

Software engineering has a complicated relationship with "correctness". We recognize the challenges of full formal rigor as well as many required properties beyond functional correctness. Although we satisfice in practice, we are still stuck in the mindset that we could reason our way to correctness, if only we had enough information. Unfortunately for our hopes of formal rigor, our expectations are shaped by unspoken knowledge that is personal, subjective, qualitative, and largely unavailable. Generative AI has introduced a new dimension to assurance: its foundation is statistical rather than formal. Traditional software engineering establishes confidence through rigorous reasoning, domain knowledge and expert judgment. In contrast, generative AI results are sophisticated predictions, "probably approximately correct". This inherently limits assurances about the results to probabilistic assertions. Further, the nuances that guide human judgment are often tacit or implicit. This knowledge casts only scant shadows into the digital record, so that critical source of knowledge is only faintly represented in AI models. We have many approaches for developing assurances that a software system does what it's expected to do, though most of them focus on code specifications rather than system requirements, let alone the system's fitness for its purpose. We have failed to develop a systematic understanding of the relative merits of the various approaches to assurance. I hope that generative AI will finally force us to tackle this. To that end, I will challenge us to think systematically about our assurance techniques, especially the role of hidden context and the challenges of AI. We need ways to make informed, reasoned choices about cost-effective combinations of approaches to developing confidence in our systems. We call ourselves software engineers. Let's act like engineers.

补充信息

↑