逃离流沙:呼吁行动
Escaping the Quicksand: A Call to Arms
浏览论文内容
中文总结 AI 辅助
本文针对计算领域技术债务带来的风险,提出结合测试、规范与证明的务实方法,呼吁社区构建语义基础设施,为AI与人类开发提供更有效的反馈循环。
中文摘要 AI 辅助
计算领域取得了惊人的成功,但累积的技术债务给我们带来了巨大的商业成本和社会风险。75年来,我们一直采用基于散文式规范的测试与调试开发方式,这种方式足以支撑行业蓬勃发展,但反馈循环昂贵且低效,让所有人都依赖于不稳定的基础。如今,AI赋能的工程技术在降低编码成本的同时放大了成功,但也通过快速增加技术债务、自动化检测其中的漏洞而放大了风险。我们该如何做得更好?长期以来,研究人员一直追求正确性的数学证明,与测试不同,证明可以覆盖所有情况,这方面也取得了巨大进展,但由于技术难度和根深蒂固的文化隔阂,证明的应用仍然困难。相反,我们主张采用务实的方法,灵活结合测试、规范和证明,为AI和人类开发提供更有效的反馈循环。最简单的方式是,开发者可以逐步将可执行的测试或神谕式的部分规范与传统的散文描述、代码和测试协同开发,这能明确设计并大幅提升测试的区分度,开发者现在就可以这么做。更好的方式是,使用支持全系列测试、基于属性的测试、符号执行和证明的规范,这能形成一系列相互交织的反馈循环,从低成本的测试到更高成本的证明,同样适用于AI和人类。然而,要真正实现实用化,还需要语义基础设施——即针对主要编程语言和其他抽象的规范与工具,我们现在大致知道如何构建这些基础设施,但它们尚未到位。我们呼吁整个社区行动起来,创建并部署这些基础设施,以构建一个建立在更坚实基础上的未来。
英文摘要
Computing has been an astonishing success - but the accumulated technical debt exposes us all to huge costs in business and societal risk. For 75 years, we've built systems to prose specifications with test-and-debug development. That works well enough for industry to thrive, but it's an expensive and ineffective feedback loop, and leaves everyone relying on shaky foundations. Now, AI-enabled engineering is amplifying the success by reducing coding costs, but also amplifies the risks, by rapidly increasing technical debt, and by automating detection of the vulnerabilities therein. How can we do better? Research has long pursued mathematical proof of correctness, which, unlike testing, can cover all cases. This too has advanced massively, but it remains hard to apply, both technically and because of a deep-seated cultural disconnect. Instead, we argue for a pragmatic approach to flexible combinations of testing, *specification*, and proof, that provides more effective feedback loops for both AI and human development. Most simply, one can incrementally co-develop executable-as-test-oracle partial specifications alongside conventional prose descriptions, code, and tests. This clarifies design and makes testing much more discriminating. Developers can and should do it today. Or, even better, one can use specifications that support the full gamut of testing, property-based testing, symbolic execution, and proof. This enables a range of intertwined feedback loops, again both for AI and humans, from cheap testing to more expensive proof. However, making it really practical needs *semantics infrastructure*: specifications and tooling for the main programming languages and other abstractions, which we now more-or-less know how to build, but which is not yet in place. We call the community to arms to create and deploy it - to enable a future built on firmer ground.
发表机构
- University of Cambridge(剑桥大学)
- Aarhus University(奥胡斯大学)
机构由 AI 辅助整理,请以论文原文为准。