AI 中文总结
该研究通过前瞻性随机实验发现,LLM理论转程序翻译中渲染器格式未产生预测的行为几何,为该领域规范格式效应设定了边界并提供了可审计研究工具。
AI 中文摘要
口头理论无法直接运行,将其转化为可执行模型需要对变量、干预措施和交互作用做出选择。我们测试了相同理论内容的呈现方式是否会系统性改变大语言模型生成的程序。在一项预先注册的前瞻性随机实验中,16个独立分配比特将32个配对渲染器槽位分配给结构化干预契约和连贯散文两种条件。两个固定的LLM快照将5个匿名理论描述转换为冻结稀疏二次语言中的320个预先授权的单次生成程序。确定性评估器测量了原子有限差分响应(H1)和混合交互响应(H2)。两个主要终点评估了跨模型匹配距离的降低和闭集同账户的可识别性,在符号线性和幅度秩流程中评估了精确随机化推断和27项预先注册的支持标准。H1和H2均返回注册裁决NOT_SUPPORTED;108项标准评估中仅有19项通过。同账户可识别性保持在随机水平附近(AUC为0.469-0.523,对照注册的0.80阈值)。一个H2匹配距离终点在符号线性流程中移动并通过了多重性校正,但对应的幅度秩结果未达到注册的效应量下限,因此不满足联合支持规则。因此,渲染器格式并未产生预先预测的均匀、家族不变、可分类的行为几何。该结果为LLM理论转程序翻译中的规范格式效应设定了具体边界,并提供了完全可审计的随机设计、数据集和软件,用于研究可执行形式化。
英文摘要
A verbal theory does not run: translating it into an executable model requires choices about variables, interventions, and interactions. We tested whether the presentation of otherwise identical theoretical content systematically changes the programs produced by large language models. In a prospective preregistered randomized experiment, 16 independent assignment bits allocated 32 paired renderer slots between a structured intervention contract and connected prose. Two pinned LLM snapshots translated five anonymous theoretical accounts, yielding 320 preauthorized single-shot programs in a frozen sparse quadratic language. A deterministic evaluator measured atomic finite-difference responses (H1) and mixed interaction responses (H2). Two primary endpoints assessed cross-model matched-distance reduction and closed-set same-account identifiability, with exact randomization inference and 27 preregistered support criteria evaluated across signed-linear and magnitude-rank pipelines. Both H1 and H2 returned the registered verdict NOT_SUPPORTED; only 19 of 108 criterion evaluations passed. Same-account identifiability remained near chance (AUC 0.469-0.523, against a registered 0.80 threshold). One H2 matched-distance endpoint moved and survived multiplicity correction in the signed-linear pipeline, but the corresponding magnitude-rank result missed the registered effect-size floor, so it did not satisfy the joint support rule. Thus renderer format did not produce the uniform, family-invariant, classifiable behavioral geometry predicted in advance. The result places a concrete boundary on specification-format effects in LLM theory-to-program translation and provides a fully auditable randomized design, datasets, and software for studying executable formalization.
Comments112 pages, 2 figures; includes supplementary material and five reproducibility annexes. Data: https://doi.org/10.5281/zenodo.21876518. Software: https://doi.org/10.5281/zenodo.21879576