十六种模型,不足两种声音:在无唯一正确答案的场景中测量集成模型的离散度
Sixteen models, fewer than two voices: measuring ensemble dispersion where no answer is uniquely correct
浏览论文内容
中文总结 AI 辅助
该研究针对无唯一正确答案的心理治疗案例,以十六种语言模型为对象,定义模型异议贡献,发现集成模型离散度由临床内容组织,模型身份为异议的可检测结构因素。
中文摘要 AI 辅助
来自十个家族的十六种语言模型,针对一个心理治疗案例,平均生成了1.69种不同表述的语义多样性,而单一模型基线为该模型自身运行得到的1.43种表述。集成模型向决策者呈现多种解读,前提是多个模型提供不同视角。对其输出的离散度从多样性和不确定性两个维度进行测量,且两种方法均依据该任务不具备的正确性标准进行验证。测量多样性是已解决的问题:Vendi Score(相似矩阵的冯·诺依曼熵的指数)是有效不同元素的数量。单一聚合结果无法说明多样性来源。我们定义了每个模型的异议贡献,即该模型与集成中其他成员平均相似度的补数——该值来自同一矩阵,而非谱指数的分解,其最大值可识别最具分歧的声音。跨模型和案例,我们将“模型身份是否能解释异议方差的非零份额”作为预先注册的假设进行测试,并刻画该测试检测到的结构。该小组构建了15个分层小场景,生成7082种表述用于分析。模型身份是异议的可检测结构因素,但常规类别仅能部分恢复它:规模差异在不同对间方向相反,家族将模型分组为仅5条两成员线,且最具分歧的声音随小组构成变化,因此浮现的异常值描述的是集成而非模型。异议未追踪案例库分层所针对的解释开放性,而是由临床内容组织,这表明集成产生的离散度是需测量的属性,而非可预设的。
英文摘要
Sixteen language models drawn from ten families produced, on average, the semantic diversity of 1.69 distinct formulations of a psychotherapeutic case, against a single-model baseline of 1.43 from one model's own runs. Ensembles place more than one reading before a decision-maker on the premise that several models supply several perspectives. Dispersion over their outputs is measured both as diversity and as uncertainty, and both traditions validate it against a correctness criterion that this task does not admit. Measuring diversity is a solved problem: the Vendi Score, the exponential of the von Neumann entropy of a similarity matrix, is an effective number of distinct elements. What a single aggregate does not say is where the diversity comes from. We define a per-model dissent contribution, the complement of a model's mean similarity to the other members of its ensemble: a magnitude from the same matrix, not a decomposition of the spectral index, whose maximum identifies the most divergent voice. Crossing model and case, we test as a preregistered hypothesis whether model identity accounts for a non-zero share of the variance in dissent, and characterise the structure that test detects. The panel formulated fifteen stratified vignettes, yielding 7,082 formulations for analysis. Model identity was a detectable structuring factor of the dissent that remained, but the usual categories recovered it only partly: scale differences pointed in opposite directions across pairs, family grouped models on only five two-member lines, and the most divergent voice changed with panel composition, so that the surfaced outlier describes the ensemble rather than the model. Dissent did not track the interpretive openness for which the case bank was stratified; it was organised by clinical content instead, leaving the dispersion an ensemble produces a property to measure rather than assume.