发表机构
Research Center for Social Computing and Interactive Robotics; Harbin Institute of Technology(社会计算与交互机器人研究中心; 哈尔滨工业大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
研究图表到代码生成,指出用参考绘图脚本监督微调存在问题,引入观察对齐监督框架,用视觉观察约束量替代潜在原始数据目标,实验证明可提升可观察值恢复。
AI 中文摘要
图表到代码生成通常在参考绘图脚本上进行监督微调,将黄金代码视为完全可观察目标。但许多图表程序含无法从渲染图像唯一恢复的潜在原始变量,监督模型重现此类不可识别量会导致幻觉和过度指定代码生成。我们引入观察对齐监督,用视觉观察约束量替代潜在原始数据目标,实验表明改进图表到代码模型需尊重图表图像可识别内容的监督目标。
英文摘要
Chart-to-code generation is commonly trained through supervised fine-tuning on reference plotting scripts, implicitly treating the gold code as a fully observable target. However, many chart programs contain latent variables that cannot be uniquely recovered from the rendered image. We identify this latent-observation mismatch in four forms across five chart types: aggregation-induced mismatch, where raw samples are reduced to box statistics or histogram bin masses; normalization-induced mismatch, where absolute scale is removed in pie charts; projection-induced mismatch, where 3D information is lost through 2D rendering; and level-set-induced mismatch, where a scalar field is observable only through selected contour lines. These mismatches introduce target ambiguity and require models to generate information unsupported by the image. We propose Observation-Aligned Supervision, which replaces latent variables with visually constrained quantities. We instantiate it using box statistics, bin weights, and wedge proportions, and study 3D scatter and contour charts through controlled experiments. Across multiple VLMs, observation-aligned supervision generally improves observable-value recovery in both-executable evaluations and mostly improves end-to-end recovery, while the contour study reveals a trade off between observation alignment and representational compactness.