arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

资本市场大语言模型可靠性评分(CM-LRS):从似是而非到可信赖

Capital Markets LLM Reliability Score (CM-LRS): From Plausible to Bankable

Prerit Ahuja

arXiv 2607.21340首次发表:更新:

AI 中文总结

研究资本市场大语言模型输出的可信赖问题,提出CM-LRS从七个维度评估工作流程输出,通过五个工作流程对四个模型与四位评判者评分,发现前沿闭源模型聚类近,差距集中在检索合成,决策有用性离散度大且评判者一致性高。

AI 中文摘要

在资本市场工作流程中,问题很少是大语言模型能否生成流畅的草稿,而是草稿是否可信赖:在对手方或监管机构面前有依据,且手头有相关文件。现有方法解决了部分差距:开放域问答基准奖励表面准确性,金融基准(FinanceBench、FinQA、ConvFinQA)推进基于文档和数值的问答,但在问答层而非从业者所维护的工作流程输出上进行评估。我们引入CM-LRS,即资本市场大语言模型可靠性评分,在工作流程输出层从七个维度评估输出:事实准确性、证据可追溯性、数值一致性、工作流程完整性、来源学科、决策有用性以及可审查性/可审计性。每个维度根据基于监管环境中审阅者使用的信号的评分标准在0至5分之间打分;总分可根据工作流程进行调整。我们在五个工作流程(债务资本市场交易条款提取、先例检索、发行人资料合成、并购交易可比推理、股权资本市场交易条款提取)上展示了CM-LRS,这些流程涉及美国证券交易委员会(SEC)的公开EDGAR文件、英国一份公开的收购公告以及虚构的合成补充材料,对四个模型与来自三个模型家族的四位独立大语言模型评判者进行评分。有三个发现。首先,前沿的闭源模型在四位评判者平均的CM-LRS上聚类在0.22分以内(Sonnet 4.6 = 4.31,Opus 4.7 = 4.30,GPT-5.5 = 4.09);所有四位评判者都将开放权重基线(Llama 3.3 70B = 3.15)排在最后。其次,差距集中在检索(2.23)和合成(2.15)上,而非提取(0.84)。第三,决策有用性在所有维度中显示出最大的跨模型离散度(在发行人资料分析中为4.0分)以及顶级的评判者间一致性(平均r = 0.52)。似是而非很容易做到。可信赖才是关键。

英文摘要

In capital-markets workflows the question is rarely whether a large language model can produce a fluent draft, but whether the draft is bankable: defensible in front of a counter-party or a regulator, with the documents in hand. Existing methods address parts of that gap: open-domain QA benchmarks reward surface accuracy, and finance benchmarks (FinanceBench, FinQA, ConvFinQA) advance document-grounded and numerical QA but evaluate at the question-answer layer rather than the workflow outputs practitioners defend. We introduce CM-LRS, a Capital Markets LLM Reliability Score, evaluating outputs at the workflow-output layer across seven dimensions: factual accuracy, evidence traceability, numerical consistency, workflow completeness, source discipline, decision usefulness, and reviewability/auditability. Each is scored 0-5 against a rubric anchored on signals reviewers in regulated settings use; the aggregate is tunable to the workflow. We demonstrate CM-LRS on five workflows (DCM transaction-terms extraction, precedent retrieval, issuer profile synthesis, M&A transaction-comparable reasoning, ECM transaction-terms extraction) over public SEC EDGAR filings, a public UK takeover release, and fictional synthetic supplements, scoring four models against four independent LLM judges spanning three model families. Three findings. First, the frontier closed-source models cluster within 0.22 points on four-judge averaged CM-LRS (Sonnet 4.6 = 4.31, Opus 4.7 = 4.30, GPT-5.5 = 4.09); all four judges place the open-weights baseline (Llama 3.3 70B = 3.15) last. Second, that gap concentrates on retrieval (2.23) and synthesis (2.15), not extraction (0.84). Third, Decision Usefulness shows the widest cross-model dispersion of any dimension (4.0 points on issuer profiling) and top-tier inter-judge agreement (mean r = 0.52). Plausibility is cheap. Bankability is the bar.

Comments23 pages. Rubrics, prompts, and demonstration tasks are publicly available

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑