Bar-JEPA:基于联合嵌入预测架构从条形图中提取数值
Bar-JEPA: Extracting Values from Bar Chart with Joint-Embedding Predictive Architecture
浏览论文内容
中文总结 AI 辅助
本文提出 Bar-JEPA,基于联合嵌入预测架构(JEPA)构建条形图数值恢复流水线,通过自监督学习提取语义特征,实现高效的条形数值恢复,效果优于端到端监督基线。
中文摘要 AI 辅助
条形图常用于数据可视化,人类易于理解,但要通过计算提取其 underlying 数据并非易事。对于基于机器学习的方法,训练图表反渲染模型通常需要带标注的真实世界数据,而标注数据的获取耗时且稀缺。当提供高语义质量的特征时,模型学习效率更高,联合嵌入预测架构(JEPA)正是为以自监督方式学习此类特征而设计的。本文提出一种针对条形图每个条形的数值恢复流水线,其中 JEPA 编码器用于生成语义丰富的隐特征;消耗这些特征的解码器模型简单且训练快速,可输出刻度和条形的坐标,用于恢复条形数值。通过将本文模型与端到端监督基线对比,自监督微调的有效性及提取特征的质量得以体现。代码、数据集和检查点已在 GitHub 上提供。
英文摘要
Bar charts are commonly used in data visualization, and while they are easily understood by humans, it is non-trivial to extract the underlying data computationally. For a machine-learning-based approach, training chart de-rendering models usually requires labeled, real-world data. Labeling data is a time consuming task, which is why annotated data is scarce. Models can learn more efficiently when provided with features of high semantic quality, which a joint-embedding predictive architecture (JEPA) is designed to learn in a self-supervised manner. We present a per-bar, numerical value recovery pipeline for bar charts, where a JEPA encoder is used to produce semantically rich latent features. The decoder model consuming these features is simple and quick to train and outputs the coordinates of ticks and bars, which can be used to recover bar values. The effectiveness of self-supervised finetuning and quality of the extracted features is evident when comparing our model to end-to-end supervised baselines. Code, datasets and checkpoints are available on \href{https://github.com/dralois/Bar-JEPA}{GitHub}.
发表机构
- Ulm University(乌尔姆大学)
机构由 AI 辅助整理,请以论文原文为准。