arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当输出分散时,认知修正会随之而来吗?面向机器集体的黑盒耦合诊断方法

When Outputs Disperse, Does Epistemic Revision Follow? A Black-Box Diagnostic for Machine Collectives

Molood Arman

arXiv 2608.03722首次发表:更新:

AI 中文总结

该研究提出黑盒耦合诊断方法,评估LLM集体输出分散度干预与认知立场修正的关联,发现不同模型的错误前提恢复效应存在显著差异,建议报告立场转变及保留前提率。

AI 中文摘要

集体智能研究将分歧视为认知多样性的证据:如果智能体表达不同观点,该群体应保留修正能力。在大型语言模型(LLM)集体中,这一代理指标可能失效:智能体可生成看似多样的论证,却保留相同结论。我们将「分散-修正耦合」操作化,即对集体在嵌入空间中的输出分散度进行可验证干预时,伴随的是认知立场的真正修正,而非保留前提的重新表述的程度。该诊断为黑盒式:仅作用于生成文本,不涉及生成模型的内部表示。我们独立测量两个通道:输出通道的一致性指数(CI)验证干预是否改变输出分散度;认知通道的逐轮立场标注,测量集体是否修正。我们提出结合元预测清晰性系统(MPCS)的CI,其在输出过度收敛时插入再分化协议(RDP),作为估计该耦合机制的可复用方法。我们评估两种配置的五智能体集体(gpt-4o-mini和gemini-2.5-flash;每个条件310对回合)。在gpt-4o-mini上,条件性异议使错误前提恢复率提升17.7个百分点(p<1e-6),而静态角色多样性损害恢复率(-8.1,p=.007)。在gemini-2.5-flash上,相同干预在可比预算下未产生增益(26.1% vs 27.1%,p=.84),尽管分散度经验证下降;两种处理效应存在差异(z=3.79,p<.001)。机制标记显示,Gemini通过框架内异议保留错误前提:94%的标记RDP后响应为重新表述而非弃权(GPT上为24%)。我们建议报告每个干预的立场转变及保留前提率,同时报告准确率。

英文摘要

Collective intelligence research treats disagreement as evidence of epistemic diversity: if agents express different views, the group should retain capacity to revise. In LLM collectives this proxy can break: agents can produce diverse-looking arguments while preserving the same conclusion. We operationalize dispersion-revision coupling: the degree to which an intervention that verifiably increases the dispersion of a collective's outputs in embedding space is accompanied by genuine revision of its epistemic stance rather than premise-preserving reformulation. The diagnostic is black-box: it operates on generated text alone and makes no claims about the internal representations of the generating models. Two channels are measured independently: an output channel, the Coherence Index (CI), verifies that the intervention changed output dispersion; an epistemic channel, per-turn stance annotation, measures whether the collective revised. We propose CI with the Meta-Predictive Clarity System (MPCS), which inserts a Re-Differentiation Protocol (RDP) when outputs over-converge, as a reusable method for estimating this coupling regime. We evaluate five-agent collectives from two configurations (gpt-4o-mini and gemini-2.5-flash; 310 paired episodes per condition). On gpt-4o-mini, conditional dissent improves false-premise recovery by +17.7 points (p<1e-6) while static persona diversity harms recovery (-8.1, p=.007). On gemini-2.5-flash, the same intervention at a comparable budget yields no gain (26.1% vs 27.1%, p=.84) despite a verified dispersion drop; the two treatment effects differ from each other (z=3.79, p<.001). Mechanism tagging shows Gemini preserves the false premise via intra-framework dissent: 94% of tagged post-RDP responses reformulate rather than concede (vs 24% on GPT). We recommend reporting per-intervention stance shift and premise-preservation rate alongside accuracy.

CommentsReviewed at Collective Intelligence 2026 (CI 2026) Conference. Revised version incorporating reviewer feedback

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑