发表机构
Globant(Globant)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文针对小型语言模型工具使用声明,提出一种廉价诊断阶梯,通过逐字复现检查和首token探针识别关键词匹配基准的误报,并用针对性SFT以极少token修复模型,提升有效调用率。
AI 中文摘要
关键词匹配基准可能为小型模型从未执行过的工具使用行为赋予分数。我们在一个架构匹配的西班牙语安全语言模型对中记录了这样一个误报案例,并提出了一套严格且廉价的诊断阶梯。一个661.6M参数的模型(约65%为代码/技术文本;无专门SFT)和一个1,109M参数的模型(以网络为主的多阶段课程;6B-token工具SFT)共享解码器、分词器和特殊token,在宽松的工具使用指标上得分几乎相同(B4:0.660对0.650)。对训练示例的逐字复现检查将它们完全区分开:600M模型在6/6个示例上生成了具有泛化参数的有效工具调用;而1B模型在所有检查点上均为0/6。一个首token探针将1B模型的失败定位为缺失先验(在<|tool_call|>上的概率为$10^{-4}$--$10^{-5}$),该先验在其以网络为主训练阶段被抹除。一种针对性的SFT方案(多样语料库,5倍学习率,2,202步,约3.3 GPU小时)用比失败阶段少三个数量级的token修复了1B模型。在所有269个语料库行上,有效生成率从0.100提升至0.959(600M:0.926)。在238个未见提示上,修复后的1B模型通过率为0.536,而600M为0.428($p = 0.004$)。嵌入漂移检查显示修复未移动触发token的绑定嵌入(bf16表中97.7%保持位相同),这意味着变化存在于周围网络中。两个模型均过度触发,很少在没有调用的情况下回答负面提示(600M为0.09,修复后的1B为0.17)。因子分析确认所有修复配置都安装了该格式,尽管多样语料库的抑制效果因种子敏感性仍是一个假设。这种廉价诊断阶梯仅需数分钟CPU时间,应作为小型模型工具使用声明的门槛。
英文摘要
Keyword-matching benchmarks can credit small models for tool use they never perform. We document such a false positive in a matched-architecture pair of Spanish security language models and propose a ladder of strict, cheap diagnostics. A 661.6M parameter model (approx. 65% code/technical text; no dedicated SFT) and a 1,109M model (web-heavy multi-phase curriculum; 6B-token tool-SFT) share decoder, tokenizer, and special tokens, scoring almost identically on lenient tool-use metrics (B4: 0.660 vs. 0.650). Verbatim-reproduction checks on training examples separate them completely: the 600M emits valid tool calls with generalized arguments on 6/6 examples; the 1B does so on 0/6 across checkpoints. A first-token probe localizes the 1B's failure to a missing prior (prob. $10^{-4}$--$10^{-5}$ on <|tool_call|>), which was erased by its web-heavy training phase. A targeted SFT recipe (diverse corpus, 5x higher learning rate, 2,202 steps, ~3.3 GPU-hours) repairs the 1B using three orders of magnitude fewer tokens than the failed phase. On all 269 corpus rows, valid emission rises from 0.100 to 0.959 (600M: 0.926). On 238 unseen prompts, the repaired 1B passes 0.536 vs. the 600M's 0.428 ($p = 0.004$). Embedding-drift checks show the repair did not move the trigger token's tied embedding (97.7% of the bf16 table remains bit-identical), meaning changes live in the surrounding network. Both models over-trigger, rarely answering negative prompts without a call (0.09 for 600M, 0.17 for repaired 1B). Factorial analyses confirm all repair configurations install the format, though suppression benefits from a diverse corpus remain a hypothesis due to seed sensitivity. This cheap diagnostic ladder costs minutes of CPU time and should gate tool-use claims on small models.
Comments24 pages, 12 tables, preprint