AI 中文总结
该研究提出推断的生成过程多样性度量,可预测38个语言模型在10个基准族上的关联失效,效果优于语义相似性,为多模型系统安全研究提供新方法。
AI 中文摘要
多样性是集体系统弹性功能中广泛存在的因素,但关键的多样性类型取决于系统的属性和失效模式,这一点对于由多个语言模型组成的系统尤为重要。不同模型即便行为与失效仍高度相关,也可能被视为独立组件。对语言模型群体的语义相似性评估显示其语义多样性有限,但该指标仅能反映观测输出含义的差异。我们提出更基础的模型多样性概念——生成过程多样性,即能够生成观测输出的过程之间的差异。基于算法信息论,我们使用原始模型输出经排列控制残差化后的归一化压缩距离,作为推断生成过程多样性的度量。在38个语言模型上,该度量识别出语义相似性遗漏的群体结构,且在10个不相交基准族的模型对间,预测了经机会校正的关联失效的跨任务变异,其效果优于语义相似性和模型对能力。跨基准的偏秩相关系数为-0.216,95%置信区间为[-0.309,-0.122],且所有10个基准上的估计值均为负。这些结果表明,生成过程多样性提升与模型对关联失效降低相关,且该关联不归因于语义相似性或能力。推断的生成过程为研究安全相关场景下多模型系统的多样性提供了新颖实用的方法。
英文摘要
Diversity is a widely observed factor in the resilient function of collective systems, yet the type of diversity that matters depends on the properties and failure modes of the system. This distinction is important for systems composed of multiple language models. Different models may be treated as independent components even when their behaviour and failures remain strongly correlated. Assessments of language-model populations using semantic similarity demonstrate limited semantic diversity, but this captures only differences in the meaning of observed outputs. We argue that a more fundamental notion of model diversity is generative-process diversity, the differences between processes capable of generating the observed outputs. Drawing from Algorithmic Information Theory, we use Normalised Compression Distance between raw model outputs, residualised against a permutation control, as a measure of inferred generative-process diversity. Across 38 language models, this measure identifies population structure missed by semantic similarity and predicts cross-task variation in chance-corrected correlated failure among model pairs across ten disjoint benchmark families, beyond semantic similarity and model-pair capability. The cross-benchmark partial rank association is $-0.216$ with a 95% interval of $[-0.309,-0.122]$, and the estimate is negative on all ten benchmarks. These results indicate that increased generative-process diversity is associated with reduced correlated failure in model pairs that is not attributable to semantic similarity or capability. Inferred generative-process diversity offers a novel and practical approach for investigating diversity of multi-model systems in safety-relevant contexts.
Comments31 pages, 13 figures