发表机构
Shanghai Jiao Tong University(上海交通大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出将自然语言到STL翻译重构为闭环反馈过程,通过反向翻译和用户修正提升准确率,实验表明强模型准确率提升至99%,弱模型提升超30个百分点。
AI 中文摘要
信号时序逻辑(STL)能够对信息物理系统进行严格的验证和控制,但编写正确的规范需要专业知识,而大多数需求持有者缺乏这种专业知识。大型语言模型可以将自然语言(NL)需求翻译为STL,然而仅靠更强的翻译器已接近准确率上限。我们认为这一上限源于任务的提出方式:一次性、开环的翻译在某种程度上是定义不明确的。自然语言具有歧义性,并且更根本的是,人们写下的内容可能并非总是其意图所在,因此目标规范并未完全包含在输入文本中。因此,我们将NL到STL的翻译重新表述为一个闭环反馈过程。每个生成的公式被翻译回自然语言供用户检查,自然语言的修正驱动修订,直到用户接受该规范。用户从不阅读或编写形式化语法。该框架基于反馈控制理论中一个熟悉的不对称性。前向路径,从歧义语言到形式逻辑,是困难且易出错的。反馈路径,从结构化STL回到语言,可以做到高度精确,而精确的反馈路径能让不精确的前向路径实现精确的闭环行为。在500条专家撰写的需求和七个大型语言模型上的实验支持了这一观点。反向翻译的解释在99.5%的情况下与专家判断一致。闭环精炼将强模型从约89%的开环准确率提升至98.0%至99.2%,并为较弱模型带来超过30个百分点的提升(例如,17.6%→48.0%)。消融实验表明,这些提升来自反馈的语义内容,而非重复尝试。专家审计和一项包含280个会话的用户研究进一步证实了该循环的可靠性。我们还识别出一个能力阈值,超过该阈值后反馈不再有帮助。
英文摘要
Signal Temporal Logic (STL) enables rigorous verification and control of cyber-physical systems, but writing correct specifications requires expertise that most requirement holders lack. Large language models can translate natural-language (NL) requirements into STL, yet stronger translators alone approach an accuracy ceiling. We argue that this ceiling stems from how the task is posed: one-shot, open-loop translation is somewhat ill-defined. Natural language is ambiguous, and, more fundamentally, what a person writes may not always be what they intend, so the target specification is not fully contained in the input text. We therefore reformulate NL-to-STL translation as a closed-loop feedback process. Each generated formula is translated back into natural language for the user to check, and natural-language corrections drive revision until the user accepts the specification. Users never read or write formal syntax. This framework rests on an asymmetry familiar from feedback control theory. The forward path, from ambiguous language to formal logic, is hard and error-prone. The feedback path, from structured STL back to language, can be made highly precise, and a precise feedback path lets an imprecise forward path achieve precise closed-loop behavior. Experiments on 500 expert-authored requirements and seven LLMs support this view. Back-translated explanations agree with expert judgments in 99.5\% of cases. Closed-loop refinement raises strong models from about 89\% open-loop accuracy to 98.0--99.2\%, and yields gains of over 30 percentage points for weaker models (e.g., 17.6\%$\rightarrow$48.0\%). Ablations show these gains come from the semantic content of the feedback rather than from repeated attempts. An expert audit and a 280-session user study further confirm the reliability of the loop. We also identify a capability threshold above which feedback no longer helps.