arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.00658cs.CV

教视觉-语言模型使用给定的尺度:用于度量物理推理的无标签等变训练

Teaching Vision-Language Models to Use the Scale They Are Given: Label-Free Equivariance Training for Metric Physical Reasoning

  • New York University(纽约大学)
  • Columbia University(哥伦比亚大学)
  • Carnegie Mellon University(卡内基梅隆大学)

机构由 AI 辅助整理,请以论文原文为准。

Kaizhen Tan, Yang Feng, Heqing Du, Siru Tao, Xin Xu, Hanzhe Hong

AI总结:

该研究针对视觉-语言模型在度量物理推理中利用尺度信息不足的问题,提出无标签等变训练方法EquiSD,通过利用物理对称性实现无监督微调,提升了模型的度量接地能力与跨尺度泛化性能。

AI中文摘要:

关于视频的度量问题要求视觉-语言模型利用提供的现实世界参考,将视觉测量值转换为物理单位。然而我们发现,当前模型仅部分利用了该尺度信息。当提示中的每个世界空间量被一个公共因子重新缩放时,视频仍然有效,正确答案恰好会按该因子变化,但模型预测仅部分调整,且准确率仍集中在所描绘物体的熟悉尺度附近。在8个视觉-语言模型中,这种反应不足的情况持续了四个数量级。当以无尺度形式询问相同物理问题时,这些模型能恢复正确的闭式尺度定律,表明主要缺陷在于度量接地而非物理机制知识。我们利用这一精确尺度关系作为监督,无需度量标注。在提供的世界空间量进行公共重新缩放时,正确的度量答案必须按相同因子变化。EquiSD利用该约束,将模型自身预测投影到尺度等变族上,并在生成的目标上微调模型。它无需真实答案,每个训练视频仅需一次模型查询。在保留的模拟视频上,EquiSD将3B模型的中位响应斜率从0.66提升至0.94,跨尺度平均相对准确率提高9.2个百分点。学习到的关系可泛化到未见过的世界尺度,且无需适配即可迁移到真实QuantiPhy视频,准确率提高6.4个百分点。这些结果表明,精确的物理对称性可作为无标签监督,用于改进视觉-语言模型的度量接地。

英文摘要:

Metric questions about video, such as the speed of a moving object, require a vision-language model to convert visual measurements into physical units using a real-world reference supplied in the prompt. We find that current models use this reference only partially. When every world-space quantity in the prompt is multiplied by a common factor, the prompt still describes the same video and the correct answer changes by exactly that factor, but the predictions of eight models change by less, and their accuracy stays concentrated near the scale that the depicted objects usually have. Asked the same physics in a scale-free form, the two models we test recover the closed-form scaling laws on most items, which indicates that the deficit lies in metric grounding and not in knowledge of the physical mechanism. Because the scaling relation is exact, it can serve as supervision without metric annotations. Equivariance Self-Distillation (EquiSD) projects a model's own prediction onto the functions that satisfy this relation and fine-tunes the model on the resulting targets, with one query per training question and no ground truth. Trained on synthetic video only, EquiSD brings a 3B model close to the exact relation on held-out simulated videos, also at scales not seen in training, and improves its accuracy across scales. Without adaptation, it also improves accuracy across scales on the QuantiPhy benchmark, where its gain reaches 93% of that obtained by supervision with exact simulator answers.

↑