arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

提示中的分析师:大语言模型金融分析中的角色、检索与记忆偏差

The Analyst in the Prompt: Role, Retrieval, and Memory Biases in LLM Financial Analysis

Ahmed Asaad, Amr Mohamed, Yang Zhang, Omneya Abdelsalam

arXiv 2609.03218首次发表:更新:

发表机构

Durham University Business School; MBZUAI; Ecole Polytechnique; HBKU(杜伦大学商学院; 穆罕默德·本·扎耶德人工智能大学; 巴黎综合理工学院; 哈迈德·本·哈利法大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

该研究针对大语言模型的金融分析场景,发现用户上下文溢出效应主要源于模型角色对证据的解读偏差,提出两种缓解策略但效果因模型而异。

AI 中文摘要

大语言模型(LLMs)越来越多地利用记忆、个人资料和角色提示等用户上下文来个性化其响应。这种个性化会影响基于证据的判断:相同的证据在不同用户上下文下可能会导致不同的结论。金融领域为研究该问题提供了高风险场景,因为决策通常依赖于对冗长复杂文档的解读。我们使用12种大语言模型的3575份美国证券交易委员会(SEC)文件对其进行测试,比较角色条件检索、中性检索以及记忆框架上下文,以区分证据选择的影响与解读的影响。我们发现,大多数用户上下文溢出效应源于模型在不同角色下对相同证据的解读方式,而非检索到不同证据。随后我们测试了两种简单的缓解策略:将相同投资者心态表述为用户个人资料而非助手角色,以及分离基于证据的输出与个性化输出。两种策略均能减少溢出效应,但均无法完全消除,且其有效性在不同模型间存在显著差异。

英文摘要

Large Language Models (LLMs) increasingly use user context such as memory, profiles, and role prompts to personalize their responses. This personalization can affect evidence-based judgment: the same evidence may lead to different conclusions under different user contexts. Finance provides a high-stakes setting to study this problem because decisions often depend on interpreting long and complex documents. We test this using 3,575 SEC filings across twelve LLMs. We compare persona-conditioned retrieval, neutral retrieval, and memory-framed context to separate the effect of evidence selection from the effect of interpretation. We find that most user-context spillover comes from how models interpret the same evidence under different roles, rather than from retrieving different evidence. We then test two simple mitigation strategies: expressing the same investor mindset as a user profile instead of an assistant role, and separating evidence-based and personalized outputs. Both reduce spillover, but neither removes it completely, and their effectiveness varies substantially across models.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑