CoVA-SFT:用于视觉抽象链的大规模数据集
CoVA-SFT: A Large-Scale Dataset for Chain of Visual Abstractions
浏览论文内容
中文总结 AI 辅助
CoVA-SFT是用于视觉抽象链的大规模数据集,含5.19万样本等,经其微调的模型在CoVA-Bench上性能优于交错式CoT基线2倍以上,凸显相关研究的开放挑战。
中文摘要 AI 辅助
思维链(CoT)推理通过将问题分解为中间步骤,极大地提升了大型语言模型(LLM)的性能。尽管CoT在语言任务中被广泛应用,但纯文本CoT会迫使模型将视觉问题序列化为生硬的文本。尽管已有处理视觉输入的架构解决方案,学界仍缺乏一个大规模、多步骤、自校正的数据集,以教导模型在解决纯文本推理问题时构建和维护内部视觉工作空间。为解决这一局限,我们推出CoVA-SFT,这是一个高度结构化的语料库,包含51900个样本,涵盖5个不同布局类别下超过222000个多模态推理步骤,以及17项复杂任务;同时推出配套基准CoVA-Bench,包含1700个保留的测试样本,覆盖相同任务以实现可复现评估。通过提供明确的理由表述、智能体渲染结果和验证循环,CoVA-SFT教导多模态语言模型交错使用文本和视觉抽象。我们通过实验验证该数据集:在CoVA-Bench上,经CoVA-SFT微调的模型平均性能优于所有交错式CoT基线2倍以上,但仍不及强大的纯文本CoT基线,凸显了未来工作的开放挑战。
英文摘要
Chain-of-thought (CoT) reasoning has dramatically improved large language models (LLMs) by allowing them to decompose problems into intermediate steps. While CoT is widely effective for linguistic tasks, text-only CoT forces models to serialize visual problems into awkward prose. Although architectural solutions exist to process visual inputs, the community lacks a massive, multi-step, self-corrected dataset to teach models how to build and maintain internal visual workspaces when solving purely textual reasoning problems. To address this limitation, we introduce CoVA-SFT, a highly structured corpus of 51.9K samples containing over 222K multimodal reasoning steps across 5 distinct layout families and 17 complex tasks, and CoVA-Bench, a companion benchmark of 1,700 held-out test samples spanning the same tasks for reproducible evaluation. By providing explicit rationale formulations, agentic renderings, and verification loops, CoVA-SFT teaches multimodal language models to interleave text and visual abstractions. We validate the dataset by demonstrating that models fine-tuned on CoVA-SFT outperform all interleaved CoT baselines by more than 2x on average on CoVA-Bench, though they still fall short of strong text-only CoT baselines, highlighting open challenges for future work.
发表机构
- University of California, Berkeley(加州大学伯克利分校)
机构由 AI 辅助整理,请以论文原文为准。