arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.24071cs.CV

可监控的图表推理智能体:基于可验证过程奖励

Monitorable Chart Reasoning Agents via Verifiable Process Rewards

Sanchit Sinha, Oana Frunza, Kashif Rasul, Aidong Zhang

首次发表
浏览论文内容

中文总结 AI 辅助

针对图表推理智能体难以审计验证的问题,提出Chart-RVR强化学习框架,将推理分解为结构、证据、推导三个可审计模块,在六个基准上达到最先进准确率,并显著提升推理的可验证性和证据基础。

中文摘要 AI 辅助

图表推理智能体越来越多地被用于在关键领域提取可操作的见解,在多个基准测试上达到了最先进的性能。然而,仅凭高基准准确率不足以支撑部署,因为利益相关者必须能够审计和验证模型如何得出其答案。现有的基于LVLM的图表智能体要么产生仅答案的预测,要么产生难以验证的自由形式推理,从而掩盖了错误是源于误读图表、提取错误数值还是计算错误。我们提出了Chart-RVR,一种用于训练可监控图表智能体的强化学习框架,该框架具有可验证的过程奖励。Chart-RVR将图表推理分解为三个可审计的模块:结构(Structure),识别图表类型;证据(Evidence),以JSON格式重建底层数据表;以及推导(Derivation),暴露计算答案的逐步轨迹。在六个域内和域外基准测试中,Chart-RVR在同等规模的LVLM中达到了最先进的准确率。除了准确率之外,我们使用三角化协议评估可监控性,该协议结合了真实标签替代指标、oracle信息增益度量以及LLM作为审计员评分过程可验证性和证据定位,表明Chart-RVR产生的推理比CoT提示、SFT和现有图表特定基线产生的推理明显更可验证和基于证据。

英文摘要

Chart reasoning agents are increasingly used to extract actionable insights in critical domains, achieving state-of-the-art performance on multiple benchmarks. Yet, high benchmark accuracy alone is insufficient for deployment, where stakeholders must be able to audit and verify how a model reaches its answer. Existing LVLM-based chart agents produce either answer-only predictions or free-form rationales that are hard to verify, obscuring whether an error arose from misreading the chart, extracting a wrong value, or miscomputing. We propose Chart-RVR, a reinforcement learning framework for training monitorable chart agents with verifiable process rewards. Chart-RVR decomposes chart reasoning into three auditable blocks: Structure, identifying the chart type; Evidence, reconstructing the underlying data table in JSON; and Derivation, exposing the stepwise trace that computes the answer. Across six in-domain and out-of-domain benchmarks, Chart-RVR attains state-of-the-art accuracy among comparable-sized LVLMs. Beyond accuracy, we assess monitorability using a triangulated protocol that combines ground-truth surrogate metrics, an oracle information-gain measure, and an LLM-as-auditor scoring Process Verifiability and Evidence Localization, showing that Chart-RVR yields rationales that are markedly more verifiable and evidence-grounded than those from CoT prompting, SFT, and existing chart-specific baselines.

发表机构

  • University of Virginia(弗吉尼亚大学)
  • Morgan Stanley(摩根士丹利)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑