BackTrend:通过反向重构评估科学弱信号预测
BackTrend: Evaluating Scientific Weak-Signal Prediction via Backward Reconstruction
- Yale NLP Lab(耶鲁大学自然语言处理实验室)
- New York University(纽约大学)
- TCS Research(塔塔咨询服务研究院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本文提出BackTrend回顾性基准,含25个AI/ML主题和66个验证弱信号,评估前沿LLM、RAG及智能体系统预测科学弱信号的能力,发现当前最佳系统F1仅10.1%,存在主题漂移等问题,额外检索证据无法弥合性能差距。
AI中文摘要:
科学弱信号是早期、低可见度的研究方向,这些方向后来成为成熟科学主题的核心,然而现有的资源,如趋势追踪、引文预测和前瞻报告,很少提供经过验证的参考集,将具体早期前兆与后来的范式联系起来。我们提出BackTrend,一个回顾性基准,在该基准中,给定一个成熟的目标主题和时间证据约束,系统必须恢复两类前兆:问题空间信号,即未被充分认识的研究问题,以及解决方案空间信号,即针对已知问题的新兴方法。BackTrend包含人工智能和机器学习领域的25个成熟目标主题和66个人工验证的弱信号,这些信号通过将每个候选信号锚定在其2019-2024年的发表频率轨迹上,从大规模文献中重构而来。我们使用语义匹配和基于覆盖率的指标评估前沿大语言模型、RAG系统和智能体研究系统。当前系统经常生成看似合理但不一致的前兆,表现出主题漂移、粒度不匹配、近似匹配和覆盖不完整;最强的系统仅达到10.1%的F1分数,而Coverage10最多覆盖参考信号的18.5%。我们的预算分析表明,额外的检索和网络搜索证据可以在适度预算内提高性能,但本身并不能弥合显著的性能差距。
英文摘要:
Scientific weak signals are early, low-visibility research directions that later become central to mature scientific topics, yet existing resources such as trend tracking, citation forecasting, and foresight reports rarely provide validated reference sets that link concrete early precursors to later paradigms. We introduce BackTrend, a retrospective benchmark in which, given a mature target topic and a temporal evidence constraint, systems must recover two types of precursors: problem-space signals, underrecognized research problems, and solution-space signals, emerging methods for known problems. BackTrend contains 25 mature target topics in artificial intelligence and machine learning and 66 human-validated weak signals, reconstructed from large-scale literature by grounding each candidate in its 2019-2024 publication-frequency trajectory. We evaluate frontier LLMs, RAG systems, and agentic research systems using semantic matching and coverage-based metrics. Current systems often generate plausible but misaligned precursors, exhibiting topic drift, granularity mismatch, near-miss matching, and incomplete coverage; the strongest system achieves only 10.1% F1, while Coverage10 reaches at most 18.5% of the reference signals. Our budget analyses show that additional retrieval and web-search evidence can improve performance up to a moderate budget, but does not by itself close the substantial performance gap.