arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

将语言模型从视觉编码器中解放出来:语义序列化作为小语言模型的感知接口

Free the Language Model From the Vision Encoder: Semantic Serialization as a Perception Interface for Small Language Models

Cong Xu, Ravi Sankar

arXiv 2609.29601首次发表:更新:

发表机构

University of South Florida(南佛罗里达大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究提出语义序列化感知接口,将视觉信息转为文本供小语言模型问答,在机器人基准上优于同规模零样本VLM,且优势随模型缩小而增强。

AI 中文摘要

端到端视觉语言模型(VLM)将视觉能力与语言模型的规模绑定在一起:随着语言模型缩小,感知和推理能力会同时下降。我们研究了一种具身场景问答(QA)接口,其中视觉从不进入语言模型。一个冻结的感知堆栈检测物体并测距;一个确定性的语义序列化器将感知状态(包括错误)编译为决策对齐的文本;一个未经修改的纯文本大语言模型(LLM)负责回答。在一个可见范围匹配、遮挡审计的校园机器人基准上,在前瞻性冻结标准下,序列化接口使用每折内域微调的检测器,其性能超过了语言模型规模同为7B的零样本VLM(0.7892对0.7462),在3B规模上差距更大(0.7673对0.6913)。预注册的解耦实验表明,该增益在释义后依然存在,将其归因于决策对齐的计算而非答案字符串泄漏,而新的判断词汇表限制了其适用范围。随着读者规模缩小到1.5B,优势增大,在0.5B时发生逆转,而一个真实值预言机定位了读者能力下限。在匹配的任务监督下,两种接口趋于一致:使用低秩适配(LoRA)微调的VLM超过了零样本系统,但仅与同等监督的文本阅读器持平(0.8441对0.8396,无统计显著差异),且两条路径仍受感知限制。报告的感知参数与VLM视觉塔的参数相当,总计算量并未减少。

英文摘要

End-to-end vision-language models (VLMs) bind visual competence to the scale of their language model: as the language model shrinks, perception and reasoning degrade together. We study an embodied scene question-answering (QA) interface in which vision never enters the language model. A frozen perception stack detects and ranges objects; a deterministic semantic serializer compiles the perceived state, errors included, into decision-aligned text; an unmodified text-only large language model (LLM) answers. On a visible-scope-matched, occlusion-audited campus-robot benchmark, under a prospectively frozen criterion, the serialized interface, using detectors fine-tuned in-domain within each fold, outperforms a zero-shot VLM whose language model has the same 7B scale (0.7892 vs 0.7462), with a larger margin at 3B (0.7673 vs 0.6913). Preregistered decoupling experiments show the gain survives paraphrase, attributing it to decision-aligned computation rather than answer-string leakage, while novel judgment vocabularies bound its scope. The advantage grows as the reader shrinks to 1.5B and reverses at 0.5B, and a ground-truth oracle locates the reader-capability floor. Under matched task supervision the interfaces converge: a VLM fine-tuned with low-rank adaptation (LoRA) overtakes the zero-shot system but only ties an equally supervised text reader (0.8441 vs 0.8396, no statistically resolved difference), and both routes remain perception-bound. Reported perception parameters are comparable to those of the VLM's vision tower, and total compute is not smaller.

Comments12 pages, 2 figures. This work has been submitted to the IEEE for possible publication. Copyright may be transferred without notice, after which this version may no longer be accessible

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑