arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

当前谁主导?用于图表转代码生成的令牌级模态仲裁

Who Leads Now? Token-Level Modality Arbitration for Chart-to-Code Generation

Qinghao Fu, Yarong Wang, Shunlei Ning, Yilin Wang, Shunwen Bai, Xinda Wang, Jiaotuan Wang, Yinan Nie, Wei Zhou

arXiv 2608.15510首次发表:更新:

AI 中文总结

提出MoCA模型,通过CAB分离视觉与编码能力并动态仲裁贡献,经两阶段训练后在三个基准测试中取得与同类模型相当的图表转代码性能。

AI 中文摘要

图表转代码生成要求模型读取图表的细粒度视觉细节,并编写可执行代码以复现该图表。现有图表转代码方法要么分别训练视觉和编码能力,要么在视觉与编码能力纠缠的图表转代码数据上进行微调,两种策略均未考虑两种能力的不同本质,也未考虑联合优化时产生的干扰。我们提出MoCA(Mixture of Cross-modal Arbitration,跨模态混合仲裁),它将两种能力分离而非融合。MoCA基于跨模态仲裁块(CAB)构建,该块将视觉分支和代码分支作为两条独立路径维护,还有一个轻量级仲裁器,在每一层及生成的令牌处仲裁两者的相对贡献。我们分两个阶段训练MoCA:首先在自蒸馏推理轨迹上进行监督预热,该轨迹将视觉理解分解为显式步骤;随后在推理过程和最终代码的奖励下进行强化学习。分析表明,仲裁器学习到的是结构化而非任意的分配,专家贡献会随令牌、层和实例系统地变化。在三个基准测试中,MoCA的性能与通用领域及图表专用模型相当。消融实验结果显示,性能提升并非仅源于更大的模型规模,而是来自互补视觉与代码分支初始化,以及通过CAB实现的输入条件仲裁的共同贡献。

英文摘要

Chart-to-code generation requires a model to read the fine-grained visual details of a chart and write executable code that reproduces it. Existing chart-to-code methods either train visual and coding abilities separately, or fine-tune on chart-to-code data with the two abilities entangled. Neither strategy accounts for the distinct nature of the two abilities or the interference that arises when they are optimized together. We propose MoCA (Mixture of Cross-modal Arbitration), which separates the two abilities rather than blending them. MoCA is built on Cross-modal Arbitration Block (CAB), which maintains a visual branch and a code branch as two distinct pathways, and a lightweight arbiter that arbitrates their relative contributions at every layer and generated token. We train MoCA in two stages: a supervised warm-up on self-distilled reasoning trajectories that decomposes visual understanding into explicit steps, followed by reinforcement learning with rewards on both the reasoning process and the final code. Analysis shows that the arbiter learns structured rather than arbitrary allocations, with expert contributions varying systematically across tokens, layers, and instances. Across three benchmarks, MoCA delivers competitive performance against general-domain and chart-specialized models. Ablation results show that the gains cannot be attributed to a larger model size alone, but instead arise from the joint contributions of complementary visual and code branch initialization and input-conditioned arbitration through CAB.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑