发表机构
Fudan University; Shanghai Innovation Institute; The Chinese University of Hong Kong; University of Oxford(复旦大学; 上海创新研究院; 香港中文大学; 牛津大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本文提出研究注意力预测(RAP)滚动基准,评估LLM智能体预测未来论文份额的能力,发现其存在目标条件化证据获取偏差,并验证了微调可提升预测性能。
AI 中文摘要
大型语言模型(LLMs)日益充当研究智能体的角色,然而,由于综述和研究想法缺乏唯一可验证的结果,评估它们追踪研究注意力转移的能力十分困难。我们引入了研究注意力预测(RAP),这是一个滚动基准,涵盖278个AI/ML领域和1,390个片段。在每个截止点,一个LLM智能体在时间受限的arXiv语料库中搜索,并预测未来六个月八个固定研究方向上的论文份额。搜索通常有帮助,但所有四个诊断模型在组合准确性上均不如精确计数的指数加权移动平均(EWMA)基线。我们识别出两个相互关联的瓶颈。在累积历史访问下,状态前向传递在所有四个诊断模型中均优于直接预测;冻结证据重放将这种逆转的一个共享组成部分与面向预测的策略检索较少的近期证据联系起来。即使有精确的历史活动,针对未来的更新仍然有限,只有GPT-5.5加上重新开放的搜索略微超过EWMA。在实现结果上进行微调,使Qwen3-4B在后期来源的保留领域上的预测斯皮尔曼相关性提高了0.105,在变化丰富的片段上也有增益。
英文摘要
Large language models (LLMs) increasingly act as research agents, yet their ability to track shifts in research attention is difficult to evaluate because reviews and research ideas lack uniquely verifiable outcomes. We introduce Research Attention Prediction (RAP), a rolling benchmark covering 278 AI/ML fields and 1,390 episodes. At each cut-off, an LLM agent searches a temporally restricted arXiv corpus and predicts the next six months' paper shares across eight frozen research directions. Search generally helps, but all four diagnostic models perform worse than an exact-count exponentially weighted moving average (EWMA) baseline in compositional accuracy. We identify two linked bottlenecks. Under cumulative-history access, State carry-forward outperforms direct Forecast for all four diagnostic models; frozen-evidence replay links a shared component of this reversal to Forecast-oriented policies retrieving a smaller share of recent evidence. Even with exact historical activity, future-specific updating remains limited, with only GPT-5.5 plus reopened Search slightly surpassing EWMA. Fine-tuning on realised outcomes improves Qwen3-4B's forecast Spearman correlation by 0.105 on held-out fields at later origins, with gains also on change-rich episodes.