arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

REChart:基于大型推理模型的推理高效图表编辑方法

REChart: Reasoning-Efficient Chart Editing with Large Reasoning Models

Yuanbang Liu, Chenxi Ruan, Yihan Hou, Qiong Luo, Wei Zeng

arXiv 2608.17414首次发表:更新:

发表机构

HKUST(GZ)(香港科技大学(广州))

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

REChart是两阶段训练框架,通过过程级监督提升图表编辑的保真度与推理效率,在两个基准上实现同规模开源模型最优性能,推理token用量降低79.0%。

AI 中文摘要

图表编辑需要基于编辑指令从参考图表图像中推断并修改可视化代码,这对多模态大型语言模型(MLLM)的细粒度视觉推理、指令遵循和可执行代码合成能力提出了挑战。具备扩展思维链(CoT)推理能力的大型推理模型(LRM)适合处理这类复杂多模态任务。但我们的初步研究发现,推理长度与图表编辑性能之间存在“倒U型”关系:过度推理常导致“过度思考”,使模型偏向幻觉的视觉细节或陷入冗余推理循环。为解决该问题,我们提出REChart,这是一个两阶段训练框架,对中间推理步骤提供过程级监督,以提升编辑保真度和推理效率。首先,我们从大型图像-指令-代码池中合成20万条高质量推理轨迹用于监督微调,采用角色专业化的Reason-Score-Refine工作流,迭代优化图表代码以提升质量。其次,我们通过强化学习优化模型,使用两种互补奖励:“保真度”奖励评估代码正确性、视觉保真度和结构一致性;“效率”奖励为每个rollout分配随机思考预算,截断推理过程,并根据最终推理片段对输出的贡献给予奖励。在ChartEdit和ChartMIMIC基准上,我们的模型在同等规模的开源模型中实现了图表编辑的最优性能,同时缓解了过度思考,且在最大思考预算为16384 token时,与基础模型相比平均推理token使用量减少了79.0%。

英文摘要

Chart editing requires inferring and modifying visualization code from a reference chart image based on an editing instruction, challenging fine-grained visual reasoning, instruction following, and executable code synthesis capabilities of MLLMs. Large reasoning models (LRMs) with extended Chain-of-Thought (CoT) reasoning are suitable for tackling such complex multimodal tasks. However, our preliminary study reveals an ``inverted-U'' relationship between reasoning length and chart-editing performance: Excessive reasoning often leads to ``overthinking,'' where models drift toward hallucinated visual details or get stuck in redundant reasoning loops. To address the gap, we introduce REChart, a two-stage training framework that provides process-level supervision over intermediate reasoning steps, improving both editing fidelity and reasoning efficiency. First, we synthesize 200k high-quality reasoning trajectories for supervised fine-tuning from a large image-instruction-code pool, using a role-specialized agentic Reason-Score-Refine workflow that iteratively refine the chart code toward higher quality. Second, we optimize the model via reinforcement learning with two complementary rewards: a \emph{fidelity} reward evaluating code correctness, visual fidelity, and structural consistency, and an \emph{efficiency} reward that assigns each rollout a random thinking budget, truncates the reasoning process, and credits the final reasoning segment according to its contribution to the output. On the ChartEdit and ChartMIMIC benchmarks, our model achieves state-of-the-art chart-editing performance among open-source models of comparable scale, while mitigating overthinking and reducing average reasoning token usage by 79.0\% under a maximum thinking budget of 16,384 tokens compared with the base model.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑