OraclePhys:面向结构力学的大语言模型微调系统框架
OraclePhys: A Systematic Framework for LLM Fine-Tuning on Structural Mechanics
浏览论文内容
中文总结 AI 辅助
本研究提出OraclePhys,一种面向结构力学的大语言模型微调系统框架,含自动评分基准、多形式监督数据集及对照训练研究,发现标签答案形式而非比特数决定微调效果,训练所得8B模型达空间结构响应任务数据精度前沿。
中文摘要 AI 辅助
对语言模型微调所内化内容通常事后诊断,本研究将其作为实验变量。OraclePhys是含三个组件的系统微调框架:OraclePhys-Bench是精确分级的结构力学基准,其有限元预言机对每个答案及反事实编辑评分,无需人工标签或LLM评判;OraclePhys-30K是字节相同结构描述的七种答案形式的监督数据集;还有跨七种形式及三种验证器角色的对照训练研究。该研究得出两项发现:其一,标签的答案形式而非其比特数因果决定微调所传授内容——排序目标会构建出分布外正向模型,未训练的基础模型处于猜测先验,标量目标最多为部分正向模型,布尔目标无检测到的效果;向量-标量的差距在第二个物理领域、第二个模型家族及改写的评估表面上依然存在。其二,书面或经分数过滤的答案会构建该能力,而优势加权分数(GRPO)虽提升奖励,但在测试的配方和预算范围内,模型在保留的物理任务上与初始状态统计等价,仅适用于路由。训练后的8B模型是首个处理空间结构响应的LLM,达到任务的数据精度前沿:优于零样本和32样本的前沿LLM,达到专家水平。标签所阐明的目标计算内容即为微调所传授内容;训练内容即为路由内容。
英文摘要
What a language model internalizes from fine-tuning is usually diagnosed after the fact. We make it an experimental variable. OraclePhys is a systematic fine-tuning framework with three components: OraclePhys-Bench, an exactly-graded structural-mechanics benchmark whose finite-element oracle scores every answer and counterfactual edit -- no human labels, no LLM judging; OraclePhys-30K, a supervision dataset of seven answer forms over byte-identical structure descriptions; and a controlled training study across the seven forms and three verifier roles. The study yields two findings. First, the label's answer form -- not its bit count -- causally determines what fine-tuning teaches: a ranking objective installs an out-of-distribution forward model where the untrained base sits at the guessing prior, a scalar objective at best a partial one, a boolean nothing detectable; the vector-scalar gulf survives a second physics domain, a second model family, and a paraphrased evaluation surface. Second, written or score-filtered answers install this capability, while advantage-weighted scores (GRPO) raise reward yet leave the model statistically equivalent to its start on held-out physics -- within the recipes and budgets tested -- sufficing only for routing. The trained 8B -- the first LLM on spatial structural response -- reaches the task's data-precision frontier: above a frontier LLM at zero- and 32-shot, at a specialist's level. What the label spells out about the target computation is what fine-tuning teaches; what you train on is what you route.
发表机构
- University of Houston(休斯顿大学)
- Tsinghua University(清华大学)
机构由 AI 辅助整理,请以论文原文为准。