AI 中文总结
研究通过消费者聊天界面评估四个视觉语言模型在视标图表上读取方向的能力,在多种推理模式、提示变体下运行,发现各模型准确率有差异,错误有特定方向,单一准确率掩盖问题,强调应多轴评估视觉语言模型。
AI 中文摘要
目标:视觉语言模型越来越多地用于通过消费者聊天界面解释医学和日常图像,但其读取方向的能力——通过翻转E视力视标测试的单一感知操作——在其实际使用的表面上的特征描述不佳。方法:我们通过其消费者聊天界面在一组锁定的七个视标图表上评估了四个生产视觉语言模型(称为Claude、GPT、GROK和Gemini):四个均匀的翻转E图表(每个主方向一个)、两个混合方向的翻转E图表和一个Snellen字母图表作为特异性对照。每个模型由多达三名操作员在两种推理模式(快速和思考)下,在两种提示变体(有和没有明确的方向解码规则)下运行。语料库包括920次可评分试验和50420次字形判断。主要结果是针对图表设计方向的字形级准确率,用威尔逊95%置信区间汇总。结果:在相同图表上,各模型的准确率从43.0%到97.0%不等,最强的模型取决于推理模式(快速模式下GPT为97.0%;思考模式下GROK为96.6%)。错误不是随机的,而是集中在特定于模型的吸引子方向上。模型内部的自一致性为96 - 100%,但准确率差异很大,将可靠性与有效性区分开来。无答案键的总体共识估计与准确率密切相关(r = 0.998)。对于一个模型,消费者界面的准确率比编程访问低25 - 27分,几乎完全在一个方向上。结论:单一的准确率数字掩盖了临床相关的、特定于方向的失败模式;在信任图像解释输出之前,应在多个轴上并在部署表面上评估视觉语言模型。
英文摘要
OBJECTIVES: Vision-language models are increasingly used to interpret medical and everyday images through consumer chat interfaces, yet their ability to read orientation - the single perceptual operation tested by the tumbling-E acuity optotype - is poorly characterized on the surfaces through which they are actually used. METHODS: We evaluated four production vision-language models (referred to as Claude, GPT, GROK, and Gemini) through their consumer chat interfaces on a locked set of seven optotype charts: four uniform tumbling-E charts (one per cardinal orientation), two mixed-orientation tumbling-E charts, and one Snellen letter chart as a specificity control. Each model was run in two reasoning modes (Fast and Thinking) under two prompt variants (with and without an explicit orientation-decoding rule) by up to three operators. The corpus comprised 920 scoreable trials and 50,420 glyph judgements. The primary outcome was glyph-level accuracy against the chart's designed orientation, summarized with Wilson 95% confidence intervals. RESULTS: Accuracy ranged from 43.0% to 97.0% across models on identical charts, and the strongest model depended on reasoning mode (GPT 97.0% in Fast mode; GROK 96.6% in Thinking mode). Errors were not random but collapsed onto a model-specific attractor direction. Models were 96-100% internally self-consistent yet ranged widely in accuracy, dissociating reliability from validity. An answer-key-free ensemble-consensus estimate tracked accuracy closely (r = 0.998). For one model, consumer-interface accuracy fell 25-27 points below programmatic access, almost entirely on a single orientation. CONCLUSIONS: A single accuracy figure conceals clinically relevant, orientation-specific failure modes; vision-language models should be evaluated along multiple axes and on the deployment surface before image-interpretation outputs are trusted.
Comments4 figures, 4 supplementary figures