发表机构
Future Principle(未来原则(公司))
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
提出基于2100多个Polymarket问题的智能体预测环境,用Brier分数奖励的GRPO训练Qwen3.5模型,使校准提升30-40%,搜索次数下降,并以5%推理成本超越Claude Opus等前沿模型。
AI 中文摘要
基于结果的强化学习可以训练语言模型预测现实世界事件,但以往的预测工作要么在训练前冻结研究上下文,要么仅在测试时部署智能体研究,因此收集证据的技能从未由奖励塑造。我们引入了一个由2100多个已解决的Polymarket问题构建的智能体预测环境、数据集和测试框架;智能体在生成时获取自己的上下文(网络搜索、页面阅读和金融时间序列,所有这些都通过分层泄漏过滤限制为每个问题截止日期之前发布的信息),我们使用Brier分数奖励下的单轮GRPO在其上训练Qwen3.5-35B-A3B(30亿激活参数)。训练改变了智能体与信息的交互方式:校准改善30-40%,搜索尝试次数从每轮3.8次降至2.25次,因为学会了证据纪律。在相同的测试框架中与四个前沿模型进行评估,训练后的策略在基于证据的预测方面也领先于所有测试的前沿模型,包括Claude Opus 4.5(软Brier 0.254对0.256,n=265),推理成本约为其5%,并且在最困难的问题(即人群本身尚未决定的问题)上优势最大。我们发布环境、数据集和每轮记录,作为时间预测智能体的可复用测试框架。
英文摘要
Outcome-based reinforcement learning can train language models to forecast real-world events, but prior forecasting work either freezes research context before training or deploys agentic research only at test time, so the skill of gathering evidence is never shaped by the reward. We introduce an agentic forecasting environment, dataset, and harness built from 2,100+ resolved Polymarket questions; the agent acquires its own context at rollout time (web search, page reading, and financial time series, all restricted by layered leak filtering to information published before each question's cutoff), and we train Qwen3.5-35B-A3B (3B active parameters) on it with single-epoch GRPO under a Brier-score reward. Training changes how the agent interacts with information: calibration improves 30-40%, and search attempts fall from 3.8 to 2.25 per rollout as evidence discipline is learned. Evaluated in an identical harness against four frontier models, the trained policy also finishes ahead of every frontier model tested at evidence-based forecasting, including Claude Opus 4.5 (soft-Brier 0.254 vs. 0.256, n=265), at about 5% of the inference cost, and its margin is widest on the hardest questions, the ones the crowd itself had not decided. We release the environment, dataset, and per-rollout records as a reusable harness for temporal forecasting agents.
CommentsAccepted at the NeurIPS 2026 Workshop on Foundation Models for Temporal Systems (FMTS). 9 pages, 4 figures. Code and data: https://github.com/afifi-yusuf/prime-forecast