前沿金融判断:智能体能否判断哪些因素会推动股票走势?
Frontier Financial Judgement: Can agents tell what might move a stock?
浏览论文内容
中文总结 AI 辅助
该研究引入前沿金融判断基准,结合多种文章创建评估项目,让智能体区分金融信息。最强智能体匹配率仅52.4%,智能体在误报率等方面有显著差异,且在多方面存在权衡,阻碍新闻流过滤在实践中的可靠部署。
中文摘要 AI 辅助
我们引入了前沿金融判断,这是一个与专业股票分析师合作开发的具有挑战性的新基准,用于评估智能体复制专家人类判断的能力。快速识别新信息、评估其影响并确定其估值影响是现实世界股票覆盖中最耗时且具有挑战性的方面之一。随着人工智能迅速增加要处理的新信息量,这变得愈发困难和重要。我们在前沿金融判断中评估的最强智能体在仅52.4%的情况下与所有专家标签匹配。我们还发现前沿智能体在估计的误报率上存在显著差异,从GPT - 5.6 Sol的约1%到Claude Sonnet 4.6的约32%。为构建该基准并使其具有现实世界代表性,我们将人工设计和标记的合成文章与实时新闻文章及历史文档相结合,创建了656个评估项目。由此产生的任务要求智能体在现实条件下区分真正新的、与估值相关的金融信息和过时的、无关紧要的或误导性的新闻。我们发现智能体在准确性、成本、误报和可靠性之间存在重大权衡,这继续阻碍新闻流过滤在实践中的可靠部署。
英文摘要
We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to assess agents' ability to replicate expert human judgements. Rapidly identifying new information, evaluating its implications and determining its valuation impact is one of the most time-consuming and challenging aspects of real-world equity coverage. This is becoming ever more difficult and important as AI rapidly increases the quantity of new information to process. The strongest agent we evaluate on Frontier Financial Judgement matches all expert labels in only 52.4% of cases. We also find significant divergence in estimated false-positive rates among frontier agents, ranging from ~1% for GPT-5.6 Sol to ~32% for Claude Sonnet 4.6. To construct the benchmark and make it representative of real-world settings, we combine human-designed and labelled synthetic articles with live news articles and historical documents, creating 656 items for assessment. The resulting task requires agents to distinguish genuinely new, valuation-relevant financial information from stale, immaterial or misleading news under realistic conditions. We find substantial trade-offs among agent accuracy, cost, false positives and reliability that continue to hinder the reliable deployment of news-flow filtering in practice.