从数字到判断:专业大语言模型智能体与欧洲上市房地产的强化学习研究
From Numbers to Judgment: Specialist LLM Agents and Reinforcement Learning for European Listed Real Estate
AI总结:
本研究以欧洲上市房地产分析为对象,提出Larix框架将分析维度映射为专业LLM智能体,结合GRPO微调Qwen3.5-9B,实现数值任务与判断任务的性能提升且可迁移至新企业与监管体系。
AI中文摘要:
本研究探讨金融分析中的本地化数值运算与整合判断是否能通过大语言模型(LLM)专业化获益。Larix将包含16个维度的欧洲上市房地产分析框架映射为8个与维度对齐的专业智能体;在固定模型、源证据、任务指令、输出格式与评分标准的前提下,我们对比了前沿LLM在整体提示与专业分解提示下的表现。针对覆盖7类监管体系的19家企业,分解提示使数值任务的整体表现提升15.8个百分点,但对判断任务的提升不稳定甚至可能降低,该模式在4种固定模板调度下保持稳定;仅使用完整框架的单智能体无法重现数值任务的提升。随后,使用与任务对齐的结构化奖励函数,通过GRPO对Qwen3.5-9B进行后训练,使开发集得分提升12.0个百分点,判断任务整体得分提升14.2个百分点,且4项子任务均有提升;该提升可迁移至未见过的企业(整体提升15.2个百分点,契约压力任务提升40.4个百分点)与未见过的监管体系(提升4.3个百分点),在3项抗记忆拆分上均实现正向迁移。因此,提示层面的分解可提升模块化数值执行,而针对性的参数调整可提升整合性金融判断。
英文摘要:
We study whether the localized numerical operations and integrative judgments of financial analysis benefit from the same form of LLM specialization. Larix maps a 16-lens European listed-real-estate analysis framework to eight lens-aligned specialists; we compare a frontier LLM under monolithic versus specialist-decomposed prompting while holding the model, source evidence, task instructions, output schema, and scoring fixed. Across 19 firms spanning seven regulatory wrappers, decomposition improves the numerical-task aggregate by 15.8 percentage points but does not reliably improve, and can reduce, performance on judgment tasks, a pattern stable across four frozen-template dispatches; a single-agent control given the complete framework does not reproduce the numerical gain. Post-training Qwen3.5-9B with GRPO using task-aligned structured rewards then raises the development-split score by 12.0 points and the judgment aggregate by 14.2 points, with gains on all four sub-ceiling tasks; the gains transfer to unseen firms (+15.2 points overall; +40.4 on covenant stress) and to unseen regulatory wrappers (+4.3), with positive transfer on all three anti-memorization splits. Prompt-level decomposition thus improves modular numerical execution, whereas targeted parameter adaptation improves integrative financial judgment.