小语言模型是否知道自己不知道什么?
Do small language models know what they don't know?
AI总结:
本研究证明词元熵在小语言模型中失效,而语义熵可有效估计不确定性,据此将不确定查询路由至更大专家模型,最高提升准确率50个百分点,且跨家族路由效果更佳。
AI中文摘要:
我们探讨了基于熵的置信度信号能否用于提高参数量少于30亿、完全在消费级硬件上运行的小语言模型(SLMs)的准确性。我们评估了七种不同的方法,包括基于词元级别的熵的提前停止、语义熵估计以及不确定性感知的路由到更大的专家模型,涉及7个模型对和5个标准NLU基准。我们的关键发现是,在小语言模型中,词元级别的熵实际上是“盲”的:在91%的数据集-模型组合中,平均词元熵接近于零,无论答案是否正确,这使得基于词元的置信度信号在此规模下不可用。我们证明,通过生成多个样本、按语义对答案进行聚类并测量分布不确定性来计算语义熵,可以恢复一个可行的置信度信号。使用语义熵将不确定的查询选择性地路由到更大的专家模型,准确率提升最高可达+50个百分点。值得注意的是,跨家族路由(例如,SmolLM 360M到Phi-3.5-mini)平均提升+22.0%,而同家族路由平均提升仅为+6.8%,这表明专家模型的质量比架构兼容性更重要。我们的结果表明,基于熵的方法在小语言模型中的价值主张不是节省计算量,而是智能的计算分配:将更多的词元花费在最需要的地方。
英文摘要:
We explore whether entropy-based confidence signals can be leveraged to improve the accuracy of Small Language Models (SLMs) with fewer than 3 billion parameters, running entirely on consumer hardware. We evaluate seven distinct approaches, including token-level entropy early stopping, semantic entropy estimation, and uncertainty-aware routing to larger expert models, across 7 model pairs and 5 standard NLU benchmarks. Our key finding is that token-level entropy is effectively blind in SLMs: in 91% of dataset-model combinations, mean token entropy is near zero regardless of answer correctness, rendering token-based confidence signals unusable at this scale. We demonstrate that semantic entropy, computed by generating multiple samples, clustering answers by meaning, and measuring distributional uncertainty, recovers a viable confidence signal. Using semantic entropy to selectively route uncertain queries to a larger expert model yields accuracy improvements of up to +50 percentage points. Notably, cross-family routing (e.g., SmolLM 360M to Phi-3.5-mini) averages +22.0% improvement compared to +6.8% for same-family routing, revealing that expert model quality matters more than architectural compatibility. Our results suggest that the value proposition for entropy-based methods in SLMs is not computational savings but intelligent compute allocation: spending more tokens where they matter most.