先验还是反馈?大语言模型在适配神经算子时依赖什么
Prior or Feedback? What an LLM Uses When Adapting Neural Operators
浏览论文内容
中文总结 AI 辅助
该研究探究LLM适配神经算子时依赖先验还是反馈,通过受控干预发现其结合任务依赖先验与反馈敏感性,且在PDE相关任务中性能优于随机搜索和贝叶斯优化。
中文摘要 AI 辅助
大语言模型(LLM)科学智能体仅依赖初始任务上下文,还是会根据实验反馈调整决策?我们在神经算子适配场景中研究该问题,其中LLM在有限试错预算下选择微调配置。在偏微分方程(PDE)族内及跨族迁移任务中,LLM在几乎所有匹配对比中都比随机搜索和贝叶斯优化取得更低的保留测试归一化均方根误差(nRMSE)。仅靠端点性能无法区分具体情况,因此我们通过受控干预验证每个归因:在观察到任何验证分数前,LLM的首个配置已在对应随机搜索池的顶部附近排名,表明存在有用的初始偏差;互补的冷启动干预显示,所选基础学习率会随PDE描述变化。一旦反馈可用,在所有测试案例中,重新分配已评估配置的验证分数会改变下一个提议,而保留值的重写则不会产生可比的聚合效应。这些干预证实,LLM的决策级动作会响应给定任务和观察到的结果,表明它结合了依赖任务的先验与对实验反馈的敏感性。
英文摘要
Do LLM scientific agents rely only on their initial task context, or do they adapt their decisions in response to experimental feedback? We study this question in neural operator adaptation, where a large language model (LLM) selects fine-tuning configurations under a limited trial budget. Across transfers within and between partial differential equation (PDE) families, the LLM achieves lower held-out test nRMSE than random search and Bayesian optimisation in nearly every matched comparison. Endpoint performance alone cannot distinguish what happens, so we verify each attribution with controlled interventions. Before observing any validation score, the LLM's first configuration already ranks near the top of the corresponding random-search pool, indicating a useful initial bias. A complementary cold-start intervention shows that the selected base learning rate shifts with the PDE description. Once feedback becomes available, reassigning validation scores among evaluated configurations changes the next proposal in every case tested, whereas a value-preserving rewrite produces no comparable aggregate effect. These interventions establish that the LLM's decision-level actions respond to the given task and observed outcomes, showing that it combines a task-dependent prior with sensitivity to experimental feedback.
发表机构
- University of Surrey(萨里大学)
- Compare the Market
机构由 AI 辅助整理,请以论文原文为准。