发表机构
Shanghai Jiao Tong University; University of Adelaide; Tsinghua University; National University of Singapore; Peking University(上海交通大学; 阿德莱德大学; 清华大学; 新加坡国立大学; 北京大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究针对金融模型无标准答案的问题,引入GAUGE基准,依据分析师实际做法评估代理构建的估值模型,通过多种方式验证,揭示高级、初级分析师及学生在模型评分上的差异,指出当前代理在模型构建与估值判断上的强弱情况,并发布相关资源。
AI 中文摘要
金融模型将公开披露信息与分析师假设相结合以进行预测和估值。虽然部分组件可机械检查,但预测、贴现率和目标价格往往有多种合理答案。现有基准倾向于根据单一专家参考对这些输出进行评分。通过使用针对同一公司独立构建的分析师模型,研究发现单参考评分中位数低,同一时期的对在隐含价格上差异大。因此引入GAUGE基准,它使用1001个供应商分类的分析师工作簿和196任务评估集,有多层观察实践包络等。通过多项研究验证,高级分析师、初级分析师和金融专业学生在失败感知分数上有差异,当前代理在模型构建上比估值判断更强。最后还发布了相关方法、数据层等。
英文摘要
Language model agents can now construct complete financial models, but it remains unclear whether they have learned the financial judgment that gives those models meaning. Existing benchmarks often anchor correctness to an expert-authored solution (e.g., numerical targets or detailed rubrics) which introduces a implicit assumption: \emph{one expert solution can serve as ground truth.} This is appropriate when finance provides a unique answer, but not when judgment is required. Evidence from professional practice challenges this assumption: When financial models built by different analysts for the same company are graded against one another, the median score is only 0.33 under standard tolerances, \emph{revealing that single-reference grading confounds professional disagreement with error.} We therefore introduce GAUGE, a benchmark that decomposes financial-model evaluation into deterministic checks, rubric judgments and numerical rules. Built from 1,001 professional valuation spanning 922 companies and all 25 GICS industry groups, GAUGE checks mechanical properties deterministically and evaluates judgment-bearing quantities against ranges in professional practice. We then validate GAUGE as a measurement instrument by testing expertise ordering, held-out professional values, and robustness to LM judgements. Our findings show that current LM agents are far better at constructing financial models than at deriving the company-specific assumptions that drive valuation, a gap that even persists after fine-tuning.
Comments55 pages, including appendices. Code, benchmark materials, and evaluation artifacts will be publicly released