arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大语言模型(LLM)能否监测经济状况?对宏观经济指标的LLM实时预报的评估

Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators

Xinyue Zhao, Ruiyi Zhang, Liqin Ye, Rui Cao, Pengtao Xie, Sudheer Chava

arXiv 2608.30110首次发表:更新:

发表机构

Georgia Institute of Technology; University of California San Diego(佐治亚理工学院; 加利福尼亚大学圣迭戈分校)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究推出抗污染基准LiveMacroEval,对比多类基准发现,具备网络搜索能力的LLM智能体对美国16项宏观经济指标的实时预报准确性与专业基准相当,展现出宏观经济实时估算潜力。

AI 中文摘要

对主要宏观经济指标进行实时预报,即官方发布前估算当前参考期内的指标值,对货币政策和金融市场至关重要,各国央行会组建由专业经济学家组成的专门团队来完成这类估算。大语言模型(LLM)智能体是该任务的有潜力候选方案,其兼具广泛的世界知识与实时网络搜索能力,且支持的查询频率高于机构实时预报。然而,评估其实时预报能力颇具挑战:GDP、CPI等主要指标被广泛报道,可能在预训练阶段被模型记忆,因此针对历史发布数据的任何评估都易受数据污染影响。为解决该问题,我们推出LiveMacroEval,这是一个实时、抗污染的基准测试,LLM智能体在每次官方发布前的预发布窗口内,为16项美国主要宏观经济指标生成每小时一次的实时预报。实时预报质量通过两项指标评估:一是针对公告窗口股票收益的LiveMacro Score,二是来自Polymarket风格模拟交易的LiveBetting Score,对比基准包括美联储地区银行的实时预报、彭博ECOS专业共识以及自回归积分滑动平均模型(auto-ARIMA)基线。在为期6个月的实验中,我们配置4款具备网络搜索能力的最先进LLM智能体,结果显示,其汇总实时预报准确性与机构和专业基准大致相当,但不同单个指标的表现差异显著。这凸显了LLM智能体作为宏观经济状况实时估算工具的潜力。

英文摘要

Nowcasting headline macroeconomic indicators, i.e., estimating an indicator's value for the current reference period before its official release, is critical for monetary policy and financial markets, and central banks devote dedicated teams of expert economists to producing such estimates. Large language model (LLM) agents are a promising candidate for this task, combining broad world knowledge with real-time web search and supporting queries at higher frequency than institutional nowcasts. Evaluating their nowcasting capability is, however, challenging: headline indicators such as GDP and CPI are widely reported and likely memorized during pretraining, so any evaluation on historical releases is vulnerable to data contamination. To address this, we introduce LiveMacroEval, a live, contamination-resistant benchmark in which LLM agents produce hourly nowcasts for sixteen major U.S. macroeconomic indicators over a pre-release window closing at each official release. Nowcast quality is assessed through a LiveMacro Score against announcement-window equity returns and a LiveBetting Score from simulated Polymarket-style trading, with Federal Reserve regional-bank nowcasts, the Bloomberg ECOS professional consensus, and an auto-ARIMA baseline as comparators. Over six months with four state-of-the-art LLM agents configured with web search, aggregate nowcast accuracy is broadly comparable to the institutional and professional benchmarks, with performance varying widely across individual indicators. This highlights LLM agents' potential as real-time estimators of macroeconomic conditions.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑