旋转了,但旋转了多少?诊断并改进VLM中的物体旋转推理
Rotated, but How Far? Diagnosing and Improving Object-Rotation Reasoning in VLMs
浏览论文内容
中文总结 AI 辅助
针对VLM在物体旋转推理中无法准确估计旋转幅度的问题,提出OR-Bench基准和RotationCue轻量解码器,通过从冻结表征中恢复粗粒度旋转信息并反馈给模型,在三个VLM上提升宏平均准确率7.9至12.6个百分点。
中文摘要 AI 辅助
视觉语言模型(VLMs)能够检测到物体在不同视角间发生了旋转,但无法可靠地判断旋转的具体幅度。我们引入了OR-Bench,一个用于物体旋转推理的细粒度基准,包含八项任务,涵盖旋转检测、旋转幅度估计和多视角旋转推理。在12个VLM中,差距十分明显:最强的模型在检测任务上接近100%的准确率,然而即使是粗略的幅度估计也接近随机水平。当被要求给出精确角度时,模型将91.8%至100%的预测集中在0°、90°和180°这三个角度上,我们将这种失败称为“典型角度坍缩”。即使在没有视觉输入的情况下,这种坍缩依然存在。表征探测显示,信息缺失只是部分原因。尽管旋转信息在更细粒度上变得难以恢复,但仍有大量粗粒度信息保留,且一个简单的线性探针在生成答案上优于模型本身。这表明VLM未能充分利用它们已经编码的旋转信息。因此,我们提出了RotationCue,一个轻量级解码器,它从VLM自身的冻结表征中恢复粗粒度旋转信息,并将其作为中间文本上下文反馈给模型。在三个VLM上,RotationCue在OR-Bench上提升了每个模型-任务组合的性能,将宏平均准确率提高了7.9至12.6个百分点,同时保持了通用能力。
英文摘要
Vision-language models (VLMs) can detect that an object has rotated across views, but cannot reliably tell by how much. We introduce OR-Bench, a fine-grained benchmark for object-rotation reasoning with eight tasks covering rotation detection, rotation magnitude estimation, and multi-view rotation reasoning. Across 12 VLMs, the gap is stark: the strongest models approach 100% accuracy on detection, yet even coarse magnitude estimation is near chance. When asked for exact angles, models place 91.8--100% of their predictions on just $0^\circ$, $90^\circ$, and $180^\circ$, a failure we term canonical-angle collapse. This collapse persists even without visual input. Representation probing shows that missing information is only part of the explanation. Although rotation information becomes less recoverable at finer granularity, substantial coarse-grained information remains, and a simple linear probe outperforms the models' generated answers. This suggests that VLMs underuse rotation information they already encode. We therefore propose RotationCue, a lightweight decoder that recovers coarse rotation information from the VLM's own frozen representations and feeds it back to the model as intermediate textual context. Across three VLMs, RotationCue improves every model--task combination on OR-Bench, raising macro-average accuracy by 7.9--12.6 points while preserving general capabilities.