arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.33257cs.LG

两个头比一个好:聚合较弱的LLM以获得更好的预测

Two Heads Are Better Than One: Aggregating Weaker LLMs for Better Forecasts

Cheng Peng, Ruixi Luo, Zhi Chen, Wei Tang

首次发表
浏览论文内容

中文总结 AI 辅助

本研究通过聚合较弱的LLM预测器,在ForecastBench上实现了超越最强单个模型的预测性能,验证了弱到强聚合的有效性。

中文摘要 AI 辅助

大型语言模型(LLMs)越来越多地被用于预测现实世界事件,但获取最强的单个预测器可能成本高昂或受到其他限制。我们研究了弱到强的预测聚合:能否将单个较弱的LLM预测器聚合起来,以超越更强的预测器?使用ForecastBench(Karger等人,2025),我们评估了16个比较组中的70个LLM预测器,每组共享超过1,000个子问题,产生了1,121个较弱模型对。在每个组内,我们通过测试Brier分数识别最强的单个模型,并评估仅由较弱预测器组成的聚合,聚合权重在独立的训练数据上学习。我们发现了弱到强改进的大量证据。学习到的线性池化在16组中的11组中识别出一个匹配或超越最强单个模型的较弱对,并在所有16组中其Brier分数差距在5%以内。我们还发现,这些改进不依赖于拥有接近最佳的组成模型,并且通常伴随着良好的校准。额外分析表明,添加更多模型并不一致地提高性能,且在实用约束下,有竞争力的较弱模型聚合仍然可用。

英文摘要

Large language models (LLMs) are increasingly used to forecast real-world events, but access to the strongest individual forecaster may be costly or otherwise constrained. We study weak-to-strong forecast aggregation: can individually weaker LLM forecasters be aggregated to outperform a stronger forecaster? Using ForecastBench (Karger et al., 2025), we evaluate 70 LLM forecasters across 16 comparison groups, each with more than 1,000 shared subquestions, yielding 1,121 weaker-model pairs. Within each group, we identify the strongest individual by test Brier score and evaluate aggregates composed exclusively of weaker forecasters, with aggregation weights learned on separate training data. We find substantial evidence of weak-to-strong improvement. Learned linear pooling identifies a weaker pair that matches or outperforms the strongest individual in 11 of 16 groups and comes within 5% of its Brier score in all 16 groups. We also find that these improvements do not rely on having a near-best constituent and are generally accompanied by good calibration. Additional analyses show that adding more models does not consistently improve performance, and competitive weaker-model aggregates also remain available under practical constraints.

↑