arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

K-OPSD:面向AEC图纸的后训练视觉语言模型的可验证在线自蒸馏

K-OPSD: Verifiable On-Policy Self-Distillation for Post-Training Vision-Language Models on AEC Drawings

Yunfei Bai, Enrico Chionna, Akash Amol, Kawaljit Singh KC, Joern Tinnemeyer

arXiv 2609.34082首次发表:更新:

发表机构

Amazon Web Services(亚马逊云服务)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出K-OPSD方法,利用可验证监督的在线自蒸馏和交叉熵损失微调Qwen3-VL,提升AEC图纸理解,在AECV-Bench上取得最高评判分数和综合准确率。

AI 中文摘要

解读建筑、工程和施工(AEC)图纸对于通用的多模态大语言模型(MLLMs)和视觉语言模型(VLMs)而言是困难的。我们提出了K-OPSD,一种用于提升AEC图纸理解能力的VLM后训练方法。基于带有可验证监督的在线自蒸馏(OPSD),我们利用模型自身的最佳N个生成结果构建一个教师模型,这些结果通过过程级验证器进行认证,并通过在暴露已验证答案的提示下进行重采样来挽救失败的提示。随后,我们通过使用交叉熵内部损失对已验证的完成结果进行训练来执行在线模型更新,该损失优于在线蒸馏所使用的有界逐词广义Jensen-Shannon散度(JSD)。使用K-OPSD,我们在AECV-Bench数据集上对Qwen3-VL模型进行微调。所得模型获得了最高的平均评判分数(0.819)和综合准确率(0.738),在与开源基线模型的竞争中取得了有竞争力的结果。该方法可迁移到域外的ArchCAD数据集,其中8B模型获益最多。我们展示了验证器套件以及持续学习和自我改进的流程,我们的结果提供了初步证据,表明验证器引导的自蒸馏是迈向更可靠的建筑图纸机器阅读的一条有前景的途径。

英文摘要

Interpreting architecture, engineering, and construction (AEC) drawings is hard for general Multimodal Large Language Models (MLLMs) and vision-language models (VLMs). We introduce K-OPSD, a VLM post-training methodology for improving AEC drawing understanding. Building on On-Policy Self-Distillation (OPSD) with verifiable supervision, we construct a teacher from the model's own best-of-N generations, certified by a process-level verifier, and rescue failed prompts by resampling under a hint that exposes the verified answer. We then perform an on-policy model update by training on verified completions with a cross-entropy inner-loss, outperforming the bounded token-wise generalized Jensen-Shannon divergence (JSD) used by on-policy distillation. Using K-OPSD, we fine-tune Qwen3-VL models on the AECV-Bench dataset. The resulting models attain the top average judge score (0.819) and combined accuracy (0.738), achieving competitive results against open-source baseline models. The recipe transfers to the out-of-domain ArchCAD dataset, where the 8B model gains most. We present the verifier suite and the continual learning and self-improving pipeline, our results provide preliminary evidence that verifier-guided self-distillation is a promising route toward more reliable machine reading of architecture drawings.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑