arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

状态依赖误差相关性塑造人工智能代理委员会中的投票阈值

State-dependent error correlations shape voting thresholds in committees of AI agents

Haifeng Li, Mo Hai

arXiv 2607.23931首次发表:更新:

AI 中文总结

研究人工智能代理委员会投票阈值,结合萨赫 - 斯蒂格利茨筛选与误差依赖性,通过齐次可交换高斯 - copula模型及对大量投票数据的估计,发现考虑依赖关系可提升相关系数、降低损失,改进投票阈值选择。

AI 中文摘要

人工智能代理委员会的聚合优势源于成员间的互补信息。经典投票假设误差独立,而语言模型误差常共现。我们将萨赫 - 斯蒂格利茨筛选与好坏情况不同的误差依赖性相结合。在齐次可交换高斯 - copula模型中,共享误差为多数投票创造了正渐近误差下限,并能改变使预期损失最小化的批准阈值。通过对28个语言模型在四个二元筛选基准上投出的174,384票进行估计,参数预测了委员会损失。结果表明,考虑依赖关系可降低损失,如全矩阵依赖模型使相关系数从0.840提升到0.967,成本敏感阈值选择下建模依赖将缩放损失从60.25降至50.77等。

英文摘要

The aggregation benefit of a committee of artificial intelligence (AI) agents comes from complementary information across members. Classical voting guarantees assume independent errors. Language-model errors often co-occur on the same cases. We combine Sah-Stiglitz screening with error dependence that can differ between good and bad cases. In a homogeneous exchangeable Gaussian-copula model, shared errors create a positive asymptotic error floor for majority voting and can change the approval threshold that minimizes expected loss. We estimate a heterogeneous extension from 174,384 votes cast by 28 language models on four binary-screening benchmarks. Parameters estimated from odd-indexed items predicted committee loss on even-indexed items. For the sampled committee composition, the full-matrix dependence model increased identity-line R^2 from 0.840 under independence to 0.967. In a design-balanced analysis, cost-sensitive threshold selection under independence reduced scaled loss from 60.25 for majority to 52.50. Modeling dependence reduced it further to 50.77, an incremental improvement of 1.73 units (95% bootstrap CI, 0.68-2.33). The overall reduction from majority was 15.73% (95% bootstrap CI, 13.41-16.75%).

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑