发表机构
Birla Institute of Technology and Science, Pilani, India(印度皮拉尼比尔拉理工学院)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究将大型语言模型嵌入DSGE模拟器,通过PPO和GRPO对比,依据模拟经济后果而非语言合理性来生成和评估政策行动。
AI 中文摘要
大型语言模型能够生成听起来合理的经济政策回应,但这并不能证明其行为与经济动态一致。我们通过将指令微调的语言模型置于六个基于Snowdrop支持的动态随机一般均衡(DSGE)模拟器中来测试这一点。在每一轮中,模型观察经济状况和经济话语的变化,选择一个有界政策行动,并接收下一个模拟状态和经济奖励。我们实现了一个通用的Python接口,用于重复回放、持久冲击、状态克隆和滚动时域模拟。这种设置产生了一个长时域信用分配问题。政策效果可能在行动采取后的几个季度才显现。PPO具有学习到的价值函数,可以通过广义优势估计将延迟奖励传播到较早的令牌。GRPO没有学习到的价值函数,而是从完整的回放回报中分配组相对优势。因此,它无法区分哪个较早的回合导致了结果;如果每次回放获得相同的回报,归一化优势为零。我们使用PPO作为主要方法,GRPO作为匹配的无评论家基线。实验还测试了方向性语义信号、奖励时域、轨迹热启动、跨模拟器迁移以及历史锚定的疫情和货币政策冲击。目标是依据模拟的经济后果而非仅凭合理的语言来评判政策行动。
英文摘要
Large language models can produce economic policy responses that sound reasonable, but this does not show that their actions are consistent with economic dynamics. We test this by placing an instruction-tuned language model inside six Snowdrop-backed dynamic stochastic general equilibrium (DSGE) simulators. At each turn, the model observes the economy and a change in economic discourse, selects a bounded policy action, and receives the next simulated state and an economic reward. We implement a common Python interface for repeated rollouts, persistent shocks, state cloning, and rolling-horizon simulation. This setting creates a long-horizon credit-assignment problem. Policy effects may appear several quarters after an action is taken. PPO has a learned value function that can propagate delayed reward to earlier tokens through generalized advantage estimation. GRPO has no learned value function and instead assigns a group-relative advantage from complete rollout returns. It therefore cannot distinguish which earlier turn caused the outcome; if every rollout receives the same return, the normalized advantage is zero. We use PPO as the primary method and GRPO as a matched critic-free baseline. The experiments also test directional semantic signals, reward horizon, trajectory warm starts, cross-simulator transfer, and historically anchored pandemic and monetary-policy shocks. The objective is to judge policy actions by their simulated economic consequences rather than by plausible language alone.