发表机构
Shenzhen University of Advanced Technology; Shenzhen Technology University(深圳先进技术大学; 深圳技术大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究针对现有视觉-文本压缩评估依赖下游任务性能的缺陷,提出解耦语义与视觉的评估框架,引入ZeroSense基准消除文本依赖,实验证实VTC质量与下游任务准确率存在显著差异。
AI 中文摘要
近期视觉-文本压缩(VTC)方法,以DeepSeek-OCR为代表,通过利用文本到图像渲染,在长上下文建模任务中实现了令人印象深刻的高token压缩率。然而,现有评估协议严重依赖下游任务性能,由于多模态大语言模型(MLLMs)具有强大的固有语言先验,这些评估指标无法准确衡量文本保留情况。本研究引入了一种新的评估框架,该框架解耦了MLLMs的能力以可靠评估VTC质量。在该框架内,我们进一步推出ZeroSense基准以确保测试样本具有低语义相关性。通过消除文本依赖,我们的基准保证评估结果纯粹反映VTC质量,不受下游模型语义推理能力的影响。在多个数据集上开展的大量实验表明,VTC质量与下游任务准确率存在显著差异,凸显了我们这种解耦评估框架的必要性。
英文摘要
Recent visual-text compression (VTC) methods, typified by DeepSeek-OCR, report impressive high token compression ratios for long-context modeling tasks by leveraging text-to-image rendering. However, existing evaluation protocols heavily rely on downstream task performance. Such evaluation metrics fail to accurately measure text preservation due to the strong inherent linguistic priors of Multimodal Large Language Models (MLLMs). In this work, we introduce a new evaluation framework that decouples MLLMs' capabilities to faithfully assess VTC quality. Within this framework, we further introduce the ZeroSense Benchmark to ensure low semantic correlation of testing samples. By eliminating textual dependencies, our benchmark guarantees that the evaluation results are purely reflective of VTC quality, unaffected by the semantic inference capabilities of downstream models. Extensive experiments across multiple datasets demonstrate that VTC quality and downstream task accuracy diverge significantly, highlighting the necessity of our decoupled evaluation framework.