AI 中文总结
本研究构建三种互补差异度,分析32个大型语言模型的行为,发现模型家族形成聚类、跨家族距离随时间减小等规律,且该方法具有标签无关性与鲁棒性。
AI 中文摘要
基准排行榜仅能总结语言模型的性能表现,却无法体现其行为与其他模型的关联,或随模型代际产生的变化。本研究通过10000个提示词组成的共享提示库,表征6个家族共32个模型的输出行为;对每个响应进行嵌入后,构建了三种互补的句子级差异度:对齐的逐提示平均距离(一种观测到的模型响应上的伪度量)、提示词分歧的PCA压缩摘要、模型内部响应几何之间无对齐的Gromov–Wasserstein差异度。我们利用这些构建,通过行为映射、家族级漂移、层次聚类、跨家族收敛及响应云离散度,研究发布日期轴上的静态组织与时间变化。在三种构建中,模型家族形成连贯聚类,其中gpt-2是全局异常值;跨家族距离随时间减小;若干近期的推理导向模型具有相对紧凑的响应云。基于逐提示最大均值差异的令牌级交叉检查与句子级平均距离高度一致(斯皮尔曼ρ=0.98),并得到相同的定性结果。我们通过测度论视角组织这些比较,明确其对齐与不变性假设;还建立了与架构无关的充分条件,将行为相似性与推理提示覆盖范围、小的超额总体对数损失及相似的有效目标分布关联起来——这是一种训练端的解释,而非对观测趋势的经验性说明。本流程无标签,且用另外三个编码器对每个响应重新编码(压缩比低至73倍)后,仍能保持秩几何、异常值及时间趋势的符号。
英文摘要
Benchmark leaderboards summarize how well a language model performs, but not how its behavior relates to that of other models or changes across generations. We characterize the output behavior of 32 models from six families using their responses to a shared bank of 10{,}000 prompts. After embedding each response, we construct three complementary sentence-level dissimilarities: an aligned mean per-prompt distance, which is a pseudometric on observed model responses; a PCA-compressed summary of prompt-wise disagreement; and an alignment-free Gromov--Wasserstein discrepancy between models' internal response geometries. We use these constructions to study static organization and temporal change on a release-date axis through behavioral maps, family-wise drift, hierarchical clustering, cross-family convergence, and response-cloud dispersion. Across the three constructions, model families form coherent clusters, with \texttt{gpt-2} as a global outlier; cross-family distances decrease over time; and several recent reasoning-oriented models have comparatively compact response clouds. A token-level cross-check based on per-prompt Maximum Mean Discrepancy closely agrees with the sentence-level mean distance (Spearman $ρ=0.98$) and recovers the same qualitative findings. We organize these comparisons through a measure-theoretic lens making their alignment and invariance assumptions explicit. We also establish an architecture-agnostic sufficient condition linking behavioral similarity to inference-prompt coverage, small excess population log-loss, and similar effective target distributions---a possible training-side account rather than an empirical explanation of the observed trends. Our pipeline is label-free, and re-encoding every response with three further encoders---down to one $73\times$ smaller---preserves the rank geometry, the outliers, and the sign of the time trend.