发表机构
Sichuan University(四川大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究微调代码语言模型中设计意图头部对CAD生成的影响,通过CADCON及错位控制等方法,发现错误头部会降低依从性,独立指标揭示指标循环性,危害具模式特异性,错误意图会误导生成。
AI 中文摘要
微调后的代码语言模型可以根据轻量级设计意图头部进行条件设定,以引导参数化CAD生成,但模型是否真正读取头部内容尚未在独立于条件设定本身的指标下进行测试,也未进行因果控制。我们研究了CADCON,它是在Qwen2.5-Coder-1.5B的LoRA微调期间添加到CadQuery样式草图拉伸程序前的一个五特征设计意图头部,并通过对生成的B-rep实体的可执行几何断言重新评分,与头部定义的正则表达式提取器没有共享代码。在三个种子和一个预先注册的{0%,40%}-前缀×{正确,错误,掩码}-头部矩阵中,我们发现:(i)在条件完成(40%前缀)时,语义错误的头部会降低模型在无条件下可以渲染的设计意图上的依从性,低于无头部基线(0.43→0.30/0.21文本/令牌)——多边形和薄几何形状;圆形和高意图在这个检查点处于基线生成水平(两个比较组中都约为0),对这种对比没有信息价值;(ii)一个错位控制——用打乱的真实头部重新训练,头部边缘相同但内容相关性被破坏——仍然有能力,但对错误头部免疫,而标准模型则不然(文本头部;在3/3个种子上交互显著,p≤4.2×10⁻³):危害需要学习到的头部→程序映射,排除边缘/机械分布转移混淆;(iii)独立指标降低了正确头部的明显好处(令牌:+0.21正则表达式→+0.02几何形状),量化了指标循环性;(iv)危害是特定于模式的——在0%前缀时,无条件基线根本无法生成有效的CAD。错误意图不是噪声:它会积极误导生成。
英文摘要
Fine-tuned code LLMs are routinely conditioned on a design-intent specification, but the correctness axis of such a signal -- a wrong intent rather than an absent one -- has not been tested, and the benefit of conditioning is usually scored with the same detector that defines the signal. We study CADCON, a five-feature design-intent header prepended to CadQuery-style programs during LoRA fine-tuning of Qwen2.5-Coder-1.5B, scoring adherence with executable geometric assertions that share no code with the header-defining extractor. On a pre-registered sample of 400 deduplicated held-out programs stratified over eleven intent profiles, at 40% prefix and three seeds, a semantically wrong header degrades adherence below the never-header-trained baseline on 3/3 seeds under both tokenizations at the program level, and on 3/3 token and 2/3 text seeds at the 298 distinct model inputs they present. Wrong-header executability is not depressed relative to that baseline. A derangement control, retrained so every program receives another program's header -- holding the header marginal fixed while destroying its correlation with the program -- saw the same programs, indices and wrong headers. Its correct-to-wrong change is -0.006/+0.016/-0.003 against 0.124/0.241/0.230 for the standard model, and the interaction is significant on 3/3 seeds (p <= 5.9e-7), so the model's sensitivity to whether the header is right or wrong requires the learned mapping. The control sits below the baseline by the same margin under a correct as under a wrong header, so we claim that sensitivity and not the below-baseline level. On features the true intent lacks, the standard model realizes a feature far more often when the wrong header names it; the control does not. Ground truth itself scores only 0.567 here, the scale on which arm levels should be read. Wrong design intent is not inert: it actively misdirects generation.
Comments47 pages, 4 figures. v5: one change in Sec. 6. The repaired-extractor re-run, previously listed as the first follow-up, has now been run under an independent pre-registration: the deficit reproduces on the repaired scoring surface under both tokenizations, bounded by the repair changing only 11.2% of scored rows. Anchor, verdict, per-row scores on both surfaces and run log are public