arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

E2A-Bench:金融图表推理中证据到行动可靠性的基准测试

E2A-Bench: Benchmarking Evidence-to-Action Reliability in Financial Chart Reasoning

Xiaoya Wang, Yutong Xu, Junjie Wang

arXiv 2609.14302首次发表:更新:

发表机构

Tsinghua University; Jinan University(清华大学; 暨南大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

针对金融图表推理,提出E2A-Bench基准,通过四个指标评估证据到行动的可靠性,发现现有幻觉分数掩盖了方向覆盖率低、验证崩溃和微调放大买卖比等问题。

AI 中文摘要

金融视觉语言模型(VLMs)能否将图表证据转化为可靠的行为建议?现有的幻觉评估大多以声明为中心;它们评估生成的陈述是否得到支持,但不评估证据是否通过推理、置信度和最终行动保持可追溯性。我们引入了E2A-Bench,一个用于金融图表推理的969个查询的基准测试,由323个HS300成分股在三种输入模态下构建,并具有确定性的OHLCV衍生证据锚点。E2A-Bench通过UCR、RCI、ECI和NDR评估接地性、推理-行动一致性、证据-置信度校准和方向覆盖率,其中NDR衡量覆盖感知的证据到行动可靠性,而非实现的交易表现。评估20个VLM揭示了三个被标量幻觉分数隐藏的失败:最低UCR模型由于仅6.4%的方向覆盖率而按NDR排名接近底部;预言机辅助验证减少了无根据的声明但可能崩溃覆盖率;金融微调将BUY:SELL比率在严格的基础-微调对中放大了4.21至4.68倍。这些结果表明,金融VLM评估应追踪完整的证据到行动链,而不是依赖单一的幻觉分数。代码和数据:此https URL

英文摘要

Can financial vision-language models (VLMs) turn chart evidence into reliable action recommendations? Existing hallucination evaluations are mostly claim-centric; they assess whether generated statements are supported, but not whether evidence remains traceable through rationale, confidence, and final action. We introduce E2A-Bench, a 969-query benchmark for financial chart reasoning, constructed from 323 HS300 constituents under three input modalities with deterministic OHLCV-derived evidence anchors. E2A-Bench evaluates grounding, reasoning-action consistency, evidence-confidence calibration, and directional coverage through UCR, RCI, ECI, and NDR, where NDR measures coverage-aware evidence-to-action reliability rather than realized trading performance. Evaluating 20 VLMs reveals three failures hidden by scalar hallucination scores: the lowest-UCR model ranks near the bottom by NDR due to only 6.4% directional coverage; oracle-aided verification reduces unsupported claims but can collapse coverage; and financial fine-tuning amplifies the BUY:SELL ratio by factors of 4.21 to 4.68 across strict base-fine-tuned pairs. These results show that financial VLM evaluation should trace the full evidence-to-action chain rather than rely on a single hallucination score. Code and data: https://github.com/wanng-ide/E2A-Bench

CommentsEMNLP Findings

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑