arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.14808physics.comp-phcs.SE

翻译中的变换:基于LLM的科学自动形式化中的两阶段结构不确定性

Transformed in Translation: Two-Stage Structural Uncertainty in LLM-Based Scientific Autoformalization

Andre Panossian

首次发表
浏览论文内容

中文总结 AI 辅助

本研究探讨基于LLM的科学自动形式化中两阶段结构不确定性,发现形式化器与递归过程均影响结果,且共同数学语言不消除翻译来源,强调模型空间构建的重要性。

中文摘要 AI 辅助

科学自动形式化将口头描述转化为可执行的数学表达,但可执行的代码并不能确定所构建的是哪个模型。我们考察了结构不确定性的两个来源:生成响应规律的形式化器,以及将该规律转化为轨迹的递归过程。在对一个公开存档的交叉实验的二次分析中,我们研究了由两个固定的语言模型形式化器从五个工程化认知描述中生成的320个响应映射,这些映射位于一个稀疏二次语法和16个随机化区块内。在保留整个区块的情况下,源描述身份识别的准确率为78.8%(随机水平为20.0%),形式化器身份识别的准确率为96.3%(随机水平为50.0%;两者p < 0.001)。程序大小是更强的单一特征族;一项预先指定的探索性比较发现,在大小之外,局部几何结构对源描述预测没有稳定的增益。在固定每个响应映射的情况下,我们随后评估了跨越33种配置和1,013,760条有限时域轨迹的五个递归族。添加反馈、投影和泄漏产生了显著不同的结果分布。关键的区别在于哪些比较得以保留:在具有已定义排序的配置对中,中位交叉递归秩一致性对于端点幅度为0.73,而对于稳定过程为0.05。因此,一种共同的数学语言并未消除翻译的来源痕迹,且在一个可观测变量下的稳健排序并不能转移到另一个可观测变量。科学自动形式化宜被研究为模型空间构建:生成的集成及其动力学嵌入都是支撑科学主张的规范的一部分。

英文摘要

Scientific autoformalization turns verbal accounts into executable mathematics, but executable code does not settle which model has been constructed. We examine two sources of structural uncertainty: the formalizer that generates a response law, and the recurrence that turns that law into trajectories. In secondary analyses of an openly archived crossed experiment, we studied 320 response maps generated by two pinned language-model formalizers from five engineered cognitive accounts within one sparse quadratic grammar and 16 randomized blocks. With whole blocks held out, source-account identity was recovered at 78.8% accuracy (chance 20.0%) and formalizer identity at 96.3% (chance 50.0%; both p < 0.001). Program size was the stronger single feature family; a pre-specified exploratory comparison found no stable source-account predictive gain from local geometry beyond size. Holding every response map fixed, we then evaluated five recurrence families spanning 33 configurations and 1,013,760 finite-horizon trajectories. Added feedback, projection and leak produced sharply different outcome distributions. The consequential distinction was which comparisons survived: median cross-recurrence rank concordance was 0.73 for endpoint magnitude but 0.05 for settling, among the configuration pairs with defined rankings. Thus a common mathematical language did not erase translation provenance, and robust ordering under one observable did not transfer to another. Scientific autoformalization is usefully studied as model-space construction: the generated ensemble and its dynamical embedding are both part of the specification supporting a scientific claim.

↑