arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2607.15718cs.SE

用于概念设计的经过验证的大语言模型驱动的合成

Verified LLM-Driven Synthesis for Concept Design

Alcino Cunha

首次发表
浏览论文内容

中文总结 AI 辅助

研究用于概念设计的大语言模型驱动合成,给出概念和反应的形式语义及验证方法,提出基于大语言模型的合成过程,探讨引导合成的方式及技术,通过评估表明不同方法各有优劣,大语言模型驱动的场景引出多数情况下可恢复预期设计。

中文摘要 AI 辅助

概念设计围绕概念构建软件系统,概念通过反应组成应用。本文首先为概念和反应给出形式语义,实现安全不变式的自动验证。接着提出基于大语言模型驱动的CEGIS风格合成过程来生成满足不变式的反应设计。研究了通过自然语言提示和正/负场景引导合成的方法,还提出大语言模型驱动的场景引出技术。评估表明仅靠不变式合成有缺陷,场景引导合成更一致,大语言模型驱动的场景引出在多数情况下能恢复预期设计,但存在局限性。

英文摘要

Concept Design structures software systems around concepts: user-facing, self-contained units of functionality with a focused purpose. Concepts are composed into applications using synchronization rules called reactions, which specify how actions in one concept trigger actions in others. This paper first gives a formal semantics for concepts and reactions, enabling automatic verification of safety invariants in applications developed with this methodology. It then presents a CEGIS-style, LLM-driven synthesis procedure for generating reaction designs that satisfy such invariants. Because many different designs can satisfy the same invariant, we study two ways of steering synthesis toward the user's intended design: natural-language prompts and positive/negative scenarios. We also propose an LLM-driven scenario elicitation technique to support early design exploration. In an evaluation on three applications and twelve design variants using one LLM configuration, invariant-only synthesis reached verified designs quickly but often produced inconsistent designs across runs, some of which were implausible, showing that invariants alone underconstrain the design task. Scenario-guided synthesis recovered intended designs more consistently than natural-language prompting, although minimal scenarios can lead to overfitting. LLM-driven scenario elicitation, where the user classifies proposed scenarios rather than authoring them from scratch, recovered the intended designs in most variants when enough scenarios were elicited, but missed behaviors and non-determinism prevented reliable coverage in all cases.

补充信息

↑