发表机构
Huazhong University of Science and Technology; Nanyang Technological University; Central South University; Harbin Engineering University(华中科技大学; 南洋理工大学; 中南大学; 哈尔滨工程大学)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对隐私友好的AIoT,提出BFG基础模型,直接从编码图像比特流进行细粒度理解,无需像素重建,并在比特流损坏下保持稳定性能。
AI 中文摘要
图像比特流细粒度理解(IBFU)旨在直接从编码的图像字节序列执行细粒度分类和语义描述生成。与传统的像素域视觉理解不同,IBFU在未将图像完全解码到像素域的情况下进行语义分析。由于推理过程中不显式重建像素级视觉内容,该范式减少了处理流程中的视觉暴露,适用于隐私友好的人工智能物联网(AIoT)应用。本文提出了比特流细粒度生成器(BFG),一种专为IBFU设计的新型基础模型。BFG由两个主要组件组成:比特流语义编码器(BSeE)和细粒度语义生成器(FSeG)。BSeE直接从编码的图像比特流中建模语义表示,无需显式像素重建,而FSeG通过自回归生成将提取的比特流语义转换为详细的自然语言描述。为了在实际AIoT场景中训练BFG并全面评估IBFU(在这些场景中图像比特流可能在传输和存储过程中遭受损坏),我们构建了一个大规模损坏比特流细粒度理解数据集(CFU-D),其中包含完整比特流和多种损坏类型及严重程度的损坏变体。实验表明,BFG在比特流损坏下仍能保持稳定的细粒度字幕生成。例如,在Stanford Dogs Caption数据集上,平均CIDEr得分仅从0.6339轻微变化到0.6077,而视觉语言模型,如Qwen-VL-Chat、BLIP-2、GLM、Gemini和GPT,则遭受严重的性能下降。本文为AIoT中隐私友好的细粒度理解提供了一种实用范式。
英文摘要
Image Bitstream Fine-grained Understanding (IBFU) aims to directly perform fine-grained classification and semantic description generation from encoded image byte sequences. In contrast to conventional pixel-domain visual understanding, IBFU conducts semantic analysis without fully decoding images into the pixel domain. Since pixel-level visual content is not explicitly reconstructed during inference, this paradigm reduces visual exposure within the processing pipeline and suits privacy-friendly Artificial Intelligence of Things (AIoT) applications. In this paper, we propose Bitstream Fine-grained Generator (BFG), a novel foundation model tailored for IBFU. BFG consists of two main components: a Bitstream Semantic Encoder (BSeE) and a Fine-grained Semantic Generator (FSeG). BSeE directly models semantic representations from encoded image bitstreams without explicit pixel reconstruction, while FSeG transforms the extracted bitstream semantics into detailed natural-language descriptions through autoregressive generation. To train BFG and comprehensively evaluate IBFU in practical AIoT scenarios, where image bitstreams may suffer corruption during transmission and storage, we construct a large-scale Corrupted-bitstream Fine-grained Understanding dataset (CFU-D), containing both intact bitstreams and corrupted variants across multiple corruption types and severity levels. Experiments show that BFG maintains stable fine-grained caption generation under bitstream corruption. For example, the performance only has slight change from 0.6339 to 0.6077 in terms of average CIDEr score on Stanford Dogs Caption dataset, while vision-language models, such as Qwen-VL-Chat, BLIP-2, GLM, Gemini, and GPT suffer severe performance decrease. This paper provides a practical paradigm for privacy-friendly fine-grained understanding in AIoT.