arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

SCP-NL2TL:用于自然语言到时间逻辑规范的带语义验证的选择性共形预测

SCP-NL2TL: Selective Conformal Prediction with Semantic Verification for Natural Language to Temporal Logic Specifications

Yixuan Wang, Licheng Luo, Yu Fu, Kaidi Xu, Yue Dong, Mingyu Cai

arXiv 2608.05439首次发表:更新:

发表机构

University of California, Riverside; City University of Hong Kong(加州大学河滨分校; 香港城市大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究提出SCP-NL2TL框架,结合选择性共形预测与语义验证,可生成自然语言到时间逻辑的规范并判断其可靠性,在三类逻辑上提升了翻译可靠性与鲁棒性,为可信自然语言接口提供了基础。

AI 中文摘要

将自然语言指令转换为机器可解释的形式化规范,能让机器人和自主系统规划、推理并形式化验证自身行为。但现有翻译模型通常会为每个输入生成一个规范,即便结果不可靠或未捕捉到用户意图,这在安全关键型应用中会产生风险。受选择性共形预测启发,我们提出一种选择性翻译框架,该框架不仅能生成形式化规范,还能确定何时可信任这些规范。可靠性由两个互补的黑盒信号评分:一是规范回译为自然语言的保真度,二是在严格语义等价下重复翻译的离散度,这两种信号在不同错误上失效,联合使用比单独使用更能清晰区分错误翻译。共形风险控制将该评分校准为接受规范或弃权(不执行)的决策,对错误规范被接受执行的比率提供无分布约束;在任何翻译尝试前,对指令嵌入的共形异常检测器会筛选出分布外输入。该框架对形式化规范语言具有通用性,在信号时间逻辑(STL)、线性时间逻辑(LTL)和几何时空逻辑(SpaTiaL)上的实验表明,其翻译可靠性、评估跨分布偏移下的鲁棒性均有所提升,且能有效实现感知不确定性的弃权(不执行)。本研究通过让AI系统识别生成的规范何时可能不可靠,为可信自然语言接口奠定了基础。

英文摘要

Translating natural language instructions into machine-interpretable formal specifications enables robots and autonomous systems to plan, reason, and formally verify their behavior. However, existing translation models typically generate a specification for every input, even when the result is unreliable or fails to capture the user's intent, creating risks in safety-critical applications. Inspired by selective conformal prediction, we propose a selective translation framework that not only generates formal specifications but also determines when they can be trusted. Reliability is scored by two complementary black-box signals, the fidelity of the specification back-translated into natural language and the dispersion of repeated translations under exact semantic equivalence, which fail on different errors and jointly separate incorrect translations more sharply than either alone. Conformal risk control calibrates this score into a decision that accepts a specification or abstains, with a distribution-free bound on the rate at which incorrect specifications are accepted for execution, and a conformal anomaly detector on instruction embeddings screens out-of-distribution inputs before any translation is attempted. The proposed framework is general across formal specification languages, with experiments on Signal Temporal Logic (STL), Linear Temporal Logic (LTL), and geometric Spatio-Temporal Logic (SpaTiaL) demonstrating improved translation reliability, robustness under the evaluated cross-tier shifts, and effective uncertainty-aware abstention. This work establishes a foundation for trustworthy natural language interfaces by enabling AI systems to recognize when generated specifications may not be reliable.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑