AI 中文总结
本文通过行为压力测试表明,预测智能体的推理并非越多越好,提出ReliabilityRoute路由策略,根据可靠性特征选择证据来源,在16个LLM年份上取得最佳平均Brier分数。
AI 中文摘要
预测智能体日益结合语言模型推理、检索、集成和校准,但尚不清楚每种行为何时应被信任。我们在ForecastBench风格的二元预测任务上研究这一问题,将选择检索、推理、服从市场先验或使用历史类比视为可观察的智能体行为,而非隐藏的实现细节。我们的核心发现是机制选择具有来源依赖性:结构化类比在某些数据生成过程中占优,而市场/群体风格和保守基线在其他情况下更优。我们引入ReliabilityRoute,一种结构干预,利用可靠性特征(如历史覆盖率、市场先验可用性、来源先验锐度、证据强度、证据分歧和预测期限)来引导预测智能体行为。一个固定的2024年拟合规则在无需硬编码来源名称决策的情况下紧密匹配手工分类,而一个前向滚动自调整规则从先前已解析的年份重新拟合阈值,并在16个后续LLM年份中在我们确定性系统中获得最佳平均Brier分数。增益适中,历史/搜索基线仍极具竞争力。因此,主要贡献是一项行为压力测试,表明更多推理并不总是更好;预测智能体应首先估计哪个证据来源应被控制,路由策略应在可审计约束下自适应,且可复现性工件可在以下网址获取:https://this URL。
英文摘要
Forecasting agents increasingly combine language-model reasoning, retrieval, ensembling, and calibration, but it remains unclear when each behavior should be trusted. We study this question on ForecastBench-style binary forecasting tasks, treating the choice to retrieve, reason, defer to a market prior, or use a historical analog as an observable agent behavior rather than a hidden implementation detail. Our central finding is that mechanism choice is source-dependent: structured analogs dominate for some data-generating processes, while market/crowd-style and conservative baselines are better for others. We introduce ReliabilityRoute, a structural intervention that steers forecasting-agent behavior using reliability features such as historical coverage, market-prior availability, source-prior sharpness, evidence strength, evidence disagreement, and horizon. A fixed 2024-fitted rule closely matches a hand taxonomy without hard-coded source-name decisions, while a walk-forward self-adjusting rule refits thresholds from previously resolved vintages and obtains the best mean Brier score among our deterministic systems across 16 later LLM vintages. The gain is modest and historical/search baselines remain highly competitive. The main contribution is therefore a behavioral stress test showing that more reasoning is not always better; forecasting agents should first estimate which evidence source deserves control, routing policies should themselves adapt under auditable constraints, and reproducibility artifacts are available at https://github.com/louiswang524/forcastagent