回溯:重放预测市场以评估大语言模型预测器
Hindcast: Replaying Prediction Markets to Evaluate LLM Forecasters
- School of Computing and Augmented Intelligence, Arizona State University(亚利桑那州立大学计算与增强智能学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
研究针对大语言模型预测器评估中答案泄露问题,提出Hindcast方法,通过设定特定过去日期评分,重放预测市场与Reddit快照,让模型读取特定时间前帖子并评分,解决泄露问题,且能随模型改进在新市场重新评估,明确检索在不同情况的作用。
AI中文摘要:
预测器通过回测进行评估,即重放已解决的问题并对系统在结果已知之前给出的概率评分。对于大语言模型,有两个渠道会将答案泄露到这个测试中。我们引入了回溯方法(Hindcast),它通过在结果在两个渠道都不存在时,将模型视为处于选定的过去日期\(t_0\)来评分,从而关闭这两个泄露渠道。回溯方法重放已解决的Polymarket预测市场与Reddit的固定快照,让模型只读取在\(t_0\)之前撰写的帖子,并根据实际发生的情况和\(t_0\)时市场自身的价格对每个预测进行评分。由于截止日期是按市场设置的且快照不变,随着模型改进,评估可以在新市场上重新运行而不会过时。关闭泄露渠道后,检索对大多数模型仍有帮助,但仅在Reddit事先讨论过该事件的地方。在存档仅包含猜测的地方,检索会产生负面影响。
英文摘要:
Forecasters are evaluated by backtesting, which replays resolved questions and grades the probability the system would have assigned before the outcome was known. For LLMs, two channels leak the answer into this test. A model that retrieves can surface reports written after the event, turning forecasting into a lookup, and each new model is trained on data closer to the event, so a question that lay in the future for last year's models sits inside this year's training data. Either way, the test grades recall while claiming to grade foresight. We introduce Hindcast, which closes both leaks by grading a model as if it stood at a chosen past date $t_0$, before the outcome existed in either channel. Hindcast replays resolved Polymarket prediction markets against a frozen snapshot of public Reddit, lets the model read only posts written before $t_0$, and scores each forecast against both what happened and the market's own price at $t_0$, itself a human forecast made from the same past information. Because the cutoff is set per market and the snapshot never changes, the evaluation re-runs on new markets as models improve, without going stale. Once the leak is closed, retrieval still helps most models, but only where Reddit discussed the event beforehand. Where the archive carried only speculation, retrieval hurts.