arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.24842cs.CLcs.AI

阅读并非使用:检索、判断与AI金融研究工作流的设计

Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows

Miao Liu, Zhizhe Liu

首次发表
浏览论文内容

中文总结 AI 辅助

该研究针对长上下文金融分析中存在的检索-整合缺口,通过实验发现工作流架构会影响LLMs将检索到的风险披露信息整合至投资判断的效果,指出AI分析师表现由模型能力与工作流架构共同决定。

中文摘要 AI 辅助

大型语言模型(LLMs)正越来越多地被部署为AI分析师,用于处理财务披露信息并支持AI辅助投资决策。然而这类系统通常仅通过其检索能力进行评估,而非评估检索到的信息是否会影响其判断。我们在长上下文金融分析中发现了检索-整合缺口:在固定核心公司信息、仅将无关上下文从2000个token调整至128000个token的情况下,我们发现风险披露对投资判断的影响降至实验噪声水平,即便直接检索仍保持准确。该模式在不同模型家族、不同判断任务及移除实际10-K文件中真实披露的实验中均得到复制;能力更强的模型会推迟但无法消除该缺口。因果记忆干预显示,压缩摘要与源文本查找共同将披露信息传递至判断中,而工作流架构决定了这种传递是否成功:分块与摘要的管道会逐出相关信息,而在决策旁进行有针对性的结构化重述则可恢复其影响。因此,AI分析师的表现由模型能力与工作流架构共同决定,基于检索的评估可能会认证那些投资判断忽略了其已明确检索到信息的系统。

英文摘要

Large language models (LLMs) are increasingly deployed as AI analysts to process financial disclosures and support AI-assisted investment decisions. Yet such systems are usually evaluated by what they can retrieve, not whether retrieved information affects their judgments. We identify a retrieval-integration gap in long-context financial analysis. Holding focal-firm information fixed and varying only unrelated context from 2,000 to 128,000 tokens, we find that a risk disclosure's influence on investment judgments falls to the experimental noise floor even as direct retrieval remains accurate. The pattern replicates across model families and judgment tasks and in experiments removing real disclosures from actual 10-K filings. More capable models postpone but do not eliminate the gap. Causal memory interventions show that compressed summaries and source-text lookup jointly transmit disclosures into judgments. Workflow architecture determines whether this transmission succeeds: chunk-and-summarize pipelines evict relevant information, whereas a targeted, structured restatement adjacent to the decision restores its influence. AI analyst performance is therefore jointly determined by model capability and workflow architecture. Retrieval-based evaluations can certify systems whose investment judgments ignore information they demonstrably retrieved.

发表机构

  • Carroll School of Management, Boston College(波士顿学院卡罗尔管理学院)
  • Columbia Business School, Columbia University(哥伦比亚大学哥伦比亚商学院)

机构由 AI 辅助整理,请以论文原文为准。

↑