arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24001cs.AI

通过推理实现多样性:利用大语言模型群体的智慧进行未来预测

Diverse by Reasoning: Harnessing the Wisdom of LLM Crowds for Future Prediction

Nirupam Chetlapalli, Yiming Liao, Min-Chun Chen, Keke Chen

首次发表
浏览论文内容

中文总结 AI 辅助

该研究提出感知行为框架构建多样化LLM群体,用K-means++聚类选代表,三模型群体预测性能优于全部25模型的投票,降本提效,凸显群体构成与代表性多样性的重要性。

中文摘要 AI 辅助

大语言模型(LLMs)正越来越多地被用于未来预测,这推动了将多个模型作为群体智慧机制的应用。然而,单纯增加群体规模并不能保证有效的多样性,因为不同的LLMs可能表现出冗余行为。我们提出一种感知行为的框架来构建多样化的LLM群体,该框架通过模型在独立开发任务上的推理轨迹来表征模型,按行为相似性对模型进行聚类,并选择代表进行集体预测。我们使用7个开发基准测试25个LLMs以进行行为多样性建模,使用2个未来预测基准测试以评估多样化群体的性能。结果表明,群体构成比群体规模更重要:基于K-means++行为聚类的三模型medoid群体在两个预测基准上的表现优于对全部25个模型的常规投票,同时减少了88%的模型调用量和约80%的推理成本。结果进一步表明,构建有效LLM群体的关键是代表性行为多样性,而非单纯最大化多样性。

英文摘要

Large language models (LLMs) are increasingly used for future prediction, motivating the use of multiple models as a wisdom-of-the-crowd mechanism. However, simply increasing crowd size does not guarantee effective diversity, as different LLMs may exhibit redundant behaviors. We propose a behavior-aware framework for constructing diverse LLM crowds. The framework characterizes models using their reasoning traces on independent development tasks, clusters models by behavioral similarity, and selects representatives for collective prediction. We evaluate 25 LLMs using seven development benchmarks for behavioral diversity modeling and two future-prediction benchmarks for evaluating diverse crowds' performance. Our results show that crowd composition can matter more than crowd size: a three-model medoid crowd based on K-means++ behavioral clustering outperforms conventional voting over all 25 models on both prediction benchmarks, while reducing model calls by 88% and inference cost by approximately 80%. The results further suggest that representative behavioral diversity, rather than simply maximizing diversity, is important for constructing effective LLM crowds

发表机构

  • University of Maryland, Baltimore County(马里兰大学巴尔的摩县分校)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑