发表机构
China Electronics Cloud Technology Co., Ltd.(中国电子云科技有限公司)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
ViCo提出面向图表复制的视觉导向编码训练框架,通过自监督预热和多步强化学习解决奖励稀疏问题,在8B模型上达到接近专有LLM的性能。
AI 中文摘要
本文解决了生成符合人类作者论文视觉标准的高质量学术图表的挑战。尽管现有AI智能体能够生成结构良好的文本和代码,但其生成的可视化往往缺乏人类设计的风格和语义保真度。采用自我反思机制的高级编码智能体表现出较差的视觉推理能力和有限的反思遵循能力,导致奖励信号稀疏,严重削弱了其强化学习(RL)效果。我们提出了ViCo,一个面向视觉导向编码的训练框架,通过迭代反思逐步使生成的图表图像与参考对齐。我们首先引入一个自监督预热阶段,该阶段通过基于一致性的剪枝增强蒙特卡洛树搜索,以合成高质量的反思轨迹,确保每个编码步骤严格遵循先前反思的结果。随后开发了一种多步强化学习算法,利用反事实基线来估计每个细化周期内反思和动作步骤的优势,从而解决奖励稀疏性问题。为了在大规模训练中实现高效奖励,我们提出了一种自动化的多层面评估框架,通过分层异构布局图结构评估图表的风格、布局和语义一致性。在三个公开基准上的实验表明,在8B模型上训练的ViCo达到了接近具有足够反思能力的专有大型语言模型的性能。
英文摘要
This paper addresses the challenge of generating high-quality academic charts that match the visual standards of human-authored papers. While existing AI agents can produce well-structured text and code, their generated visualizations often lack the stylistic and semantic fidelity of human designs. Advanced coding agents that employ self-reflection mechanisms exhibit poor visual reasoning and limited reflection following, resulting in sparse reward signals that severely undermine their reinforcement learning (RL). We propose ViCo, a training framework for visual-oriented coding that employs iterative reflections to align generated chart images progressively with the reference. We first introduce a self-supervised warm-up stage, which augments Monte Carlo Tree Search with consistency-based pruning to synthesize high-quality reflection trajectories, ensuring that each coding step strictly follows the outcomes of prior reflections. A multi-step RL algorithm is then developed, using counterfactual baselines to estimate advantage for reflection and action steps within each refinement cycle, thereby addressing the reward sparsity. To enable efficient reward in massive training, we propose an automatic, multifaceted evaluation framework that assesses charts' style, layout, and semantic consistency via a hierarchical heterogeneous layout graph structure. Experiments on three public benchmarks demonstrate that ViCo, trained on an 8B model, achieves performance close to proprietary LLMs with adequate reflection capabilities.
Comments29 pages, 7 figures. To appear in the Proceedings of EMNLP 2026 Findings