发表机构
Georgia Institute of Technology; Florida Institute of Technology(佐治亚理工学院; 佛罗里达理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
通过OperatorCLIP对比实验,发现文本条件神经代理模型的误差降低不能证明其理解文本,恒定条件控制更优,提示干预无可靠语义排序,强调路径控制的必要性。
AI 中文摘要
文本条件神经代理模型的误差降低本身并不能证明模型使用了文本的含义。我们使用OperatorCLIP研究这一归因问题,比较了无条件的FNO、恒定句子的FiLM控制以及使用对比对齐训练的固定任务描述。三种子实验覆盖了Darcy2D、ShallowWater2D和三维可压缩纳维-斯托克斯(CNS3D)任务。恒定条件控制在两个二维任务上具有更低的平均测试误差。相对于该控制,任务文本加对齐在ShallowWater2D和CNS3D上具有相似的平均值,而在Darcy2D上具有更高的平均值;这些描述性比较存在显著的种子不确定性。后一种比较同时改变了提示内容和损失函数,因此无法隔离任一效应。文本编码器从头开始训练,每个条件模型在训练期间仅看到一种描述。在这种设置下,成对InfoNCE无法识别匹配对,并具有最小$\log B$。提示干预未显示可靠的语义排序。这一方法论警示证明了路径控制的必要性;它既未确立编码器的语义能力,也未测试文本在不同物理情境下的有效性。
英文摘要
Lower error from a text-conditioned neural surrogate does not, by itself, show that the model uses the meaning of the text. We examine this attribution problem with OperatorCLIP, comparing an unconditioned FNO, a constant-sentence FiLM control, and a fixed task description trained with contrastive alignment. Three-seed experiments cover Darcy2D, ShallowWater2D, and three-dimensional compressible Navier-Stokes (CNS3D). Constant conditioning has lower mean test error on both 2D tasks. Relative to this control, task text plus alignment has a similar mean on ShallowWater2D and CNS3D and a higher mean on Darcy2D; these descriptive comparisons have substantial seed uncertainty. The latter comparison changes both prompt content and loss, so it isolates neither effect. The text encoder is trained from scratch, and each conditioned model sees only one description during training. In this regime, pairwise InfoNCE cannot identify matched pairs and has minimum $\log B$. Prompt interventions show no reliable semantic ordering. This methodological caution demonstrates why pathway controls are needed; it neither establishes semantic competence of the encoder nor tests the effectiveness of text under varying physical context.
Comments8 pages, 2 figures, 2 tables. Accepted to NeurIPS 2026 Workshop on Representation for the Physical Sciences