AI 中文总结
本文推出SocietyBench基准,通过构建反事实社会世界,测试LLMs预测社会事件的概率校准与时间准确性,发现前沿LLMs表现远逊于基准,智能体框架未提升性能,强调多事件评估的必要性。
AI 中文摘要
大型语言模型(LLMs)以及基于它们构建的智能体,如今在能否完成任务(修复漏洞、操控浏览器、操作图形用户界面)方面接受了大量基准测试。而一项互补的社会能力——即模型理解和预测真实社会事件发展方式的能力,却几乎未被测量。我们推出SocietyBench,这是一个端到端基准,它接收一行事件主题,收集来自五个平台的网络新闻和社交媒体帖子,将其提炼为按日期索引的时间线,该时间线将事实事件与公众意见层分开,随后将时间线上的每个截止日期转化为经过审核的预测问题库。这些问题在两个正交的100分轴上评分:概率校准和时间准确性。在任何模型查看时间线之前,会有一个三阶段程序替换每个命名实体,并将每个日期按每个事件的常数偏移,将真实的事件序列转化为反事实社会世界——其结构与实际发生的情况相同,但去除了模型可匹配预训练记忆的表面标签。在五个不同的事件以及中文和英文版本的125个预测点上,六个前沿LLMs中最强的一个仅达到75.0分(满分100),而一个简单基准的得分为50分。两个轴是分离的:模型可能校准能力强但时间能力弱,反之亦然。基于共享基础模型构建的三个智能体框架未能提升该基础模型的性能,两个无模型启发式方法的表现落后于所有LLMs。单个事件的轴间差距达到21.4分,这是我们主张对多个事件而非单个事件进行评估的主要理由。所有匿名时间线、问题库、真实值和评分代码均已发布。
英文摘要
Large language models (LLMs), and the agents built on top of them, are now benchmarked heavily on whether they can finish a task -- fix a bug, drive a browser, operate a GUI. A complementary social ability, namely how well a model understands and forecasts the way real social events unfold, has barely been measured. We introduce SocietyBench, an end-to-end benchmark that takes a one-line event topic, collects Web news and social-media posts across five platforms, distills them into a date-indexed timeline that keeps factual events and a public-opinion layer separate, and then turns every cutoff date on that timeline into an audited bank of forecasting questions. Questions are scored on two orthogonal 100-point axes: probability calibration and temporal accuracy. Before any model sees a timeline, a three-phase procedure replaces every named entity and shifts every date by a per-event constant, turning a real arc into a counterfactual social world -- structurally identical to what happened, but stripped of the surface labels a model could match against pre-training memory. On five heterogeneous events and 125 prediction points in Chinese and English editions, the strongest of six frontier LLMs reaches only 75.0 out of 100, against a trivial anchor of 50. The two axes come apart: a model can be calibration-strong but time-weak, or the reverse. Three agent frameworks built on a shared base model fail to improve on that base, and two model-free heuristics trail every LLM. Per-event gaps reach 21.4 points on a single axis, which is our main argument for evaluating on several events rather than one. All anonymized timelines, question banks, ground truth, and scoring code are released.
CommentsProject page: https://co-minder.github.io/Societybench