arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.01873cs.AIcs.MA

认知 Sybil 抗性:在不增加证据的情况下倍增 AI 智能体

Epistemic Sybil Resistance: Multiplying AI Agents Without Multiplying Evidence

Marc Bara

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对多智能体 AI 系统的认知 Sybil 问题,提出应追踪证据谱系与依赖性而非智能体/报告数量或相似性,通过实验验证了相关预测并设计了对应聚合器。

中文摘要 AI 辅助

多智能体 AI 系统通过生成智能体并综合报告来提升推理能力,但另一个智能体并非另一条观测:看似独立的报告可能源自同一证据,而真正独立的证据也可能产生几乎相同的报告。我们将此形式化为认知 Sybil 问题。当报告 Z 相对于报告 R 满足 I(Θ; Z | R) = 0 时,Z 是认知 Sybil 扩展。仅基于报告的聚合器通常无法区分复制与独立确证:在存在未观测到的谱系时,相同报告可对应不同的后验概率。高斯共同根源模型表明,共同谱系并不意味着完全冗余;重复提取会向源级上限增加信息,而共享基础模型在独立智能体间引发的相关提取误差会进一步降低该上限。我们在合成证据文档上测试了超过 20000 次受控 LLM 智能体报告与提取调用的预测:固定一个证据根源,报告数量从 1 增至 32 时,朴素后验覆盖率从 0.940 降至 0.263;固定报告数量,证据根源数量从 1 增至 16 时,差距缩小,且在 k=16 时聚合器统计上无差异。智能体的复制提取误差呈相关性(gamma_cal = 0.719,样本外估计),相关提取聚合器相应恢复了校准。受控操作将表征相似性与证据谱系分离:它使报告空间去重机制的平均推断聚类数变化 1.425(95% 置信区间 [1.363, 1.485]),而真实谱系的四倍变化仅使其变化 0.040([-0.045, 0.120])。因此,集体推理应追踪证据谱系与依赖性,而非智能体或报告的数量或相似性。

英文摘要

Multi-agent AI systems improve inference by spawning agents and synthesizing reports. But another agent is not another observation: apparently independent reports may descend from the same evidence, and genuinely independent evidence can produce nearly identical reports. We formalize this as an epistemic Sybil problem. A report Z is an epistemic Sybil extension relative to reports R when I(Theta; Z | R) = 0. No report-only aggregator can generally distinguish replication from independent corroboration: identical reports can warrant different posteriors under unobserved ancestry. A Gaussian shared-root model shows common ancestry does not imply complete redundancy. Repeated extraction adds information toward a source-level ceiling, and correlated extraction errors, which a shared base model can induce among independent agents, lower that ceiling further. We test these predictions with more than 20,000 controlled LLM-agent report and extraction calls on synthetic evidentiary documents. Holding one evidence root fixed while report multiplicity rises from 1 to 32 collapses naive posterior coverage from 0.940 to 0.263. Holding report count fixed while evidence-root multiplicity rises from 1 to 16 closes the gap, and the aggregators are statistically indistinguishable at k = 16. The agent's replicate extraction errors are correlated (gamma_cal = 0.719, estimated out of sample), and a correlated-extraction aggregator restores calibration accordingly. A controlled manipulation isolates representation similarity from evidential ancestry. It changes a report-space deduplication mechanism's mean inferred cluster count by 1.425 (95% CI [1.363, 1.485]), whereas a fourfold change in true ancestry changes it by only 0.040 ([-0.045, 0.120]). Collective inference should therefore track evidential ancestry and dependence, not agent or report multiplicity or similarity.

发表机构

  • Universitat Oberta de Catalunya (UOC)(加泰罗尼亚开放大学(UOC))

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑