arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

大型语言模型是好的金融用户模拟器吗?一项初步研究

Are LLMs Good Financial User Simulators? Multi-view Investor Logic Alignment (MILA)

Jiajie He, Jiangyuan Hong, Xintong Chen, Dongling Ni, Wenjin Liu

arXiv 2609.15727首次发表:更新:

发表机构

Hithink Research; Nanyang Technological University; McMaster University(恒生研究院; 南洋理工大学; 麦克马斯特大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究通过120名志愿者的模拟交易实验,初步评估了大型语言模型作为金融用户模拟器的能力,发现市场信息可提升操作与股票预测,但交易规模预测困难,且存在行为压缩现象。

AI 中文摘要

大型语言模型(LLMs)越来越多地被用作用户模拟器,但它们再现不断演变的个人金融决策的能力仍不清楚。我们在一个受控的模拟交易环境中进行了一项初步研究,涉及120名志愿者。参与者在实时市场条件下使用不可赎回的虚拟资金;未访问任何真实经纪账户、真实资金头寸或真实交易记录。仅根据预测截止日期前可获取的信息,模拟器预测参与者下一个交易日的操作、交易的证券以及交易数量。我们评估了时间对齐的滚动预测,并比较了有无时点市场信息的设置。在受控消融实验中,市场背景信息改善了操作和股票代码预测,而交易规模预测仍然困难。我们还观察到系统性的行为压缩:模型过度产生持有操作,低估卖出决策,并简化多证券交易。这些结果提供了初步的经验特征描述,并推动了对个体、时间和投资组合层面行为保真度的更大规模评估。

英文摘要

Large language models (LLMs) are increasingly used as user simulators, yet it remains unclear whether their predictions faithfully reproduce the evolving decisions of individual users. We investigate this question in a controlled longitudinal paper-trading study with 80 participants, where user interactions, simulated transactions, virtual portfolio states, and point-in-time market information are aligned under a rolling next-day prediction protocol. We evaluate behavioral fidelity hierarchically, from trade occurrence to action structure, asset selection, and downstream portfolio consequences. Across 1,239 aligned user-days, no evaluated LLM reliably outperforms a simple recent-activity persistence baseline for predicting whether a user trades. Fidelity further deteriorates at finer levels: models struggle to recover buy--sell structure and traded assets, and similar activity-level predictions can lead to substantially different portfolio trajectories. Controlled evidence ablations show that recent trading history strongly governs activity prediction, whereas asset selection is substantially more sensitive to the available evidence. An observational analysis further finds that intensified ticker-specific research predicts imminent trading, but diagnostic tests do not support a causal interpretation. These findings suggest that current LLMs capture useful short-term behavioral regularities without yet recovering a stable individual decision mechanism.

CommentsThe complete version will be open and the paper is under review in AAAI

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑