视觉语言模型能否从行人视角图像可靠评估人行道无障碍属性?
Can VLMs Reliably Assess Sidewalk Accessibility Attributes from Pedestrian-Level Imagery?
- Senseable City Lab(可感知城市实验室)
- Massachusetts Institute of Technology(麻省理工学院)
机构由 AI 辅助整理,请以论文原文为准。
AI总结:
本研究测试四种视觉语言模型从行人视角图像评估人行道无障碍属性的可靠性,引入保形预测校准不确定性,发现有效宽度评估较有效,但多数属性精度不足,并公开了带标注图像数据集。
AI中文摘要:
城市无障碍环境的重要组成部分,特别是对于轮椅使用者和行动不便人群,是人行道是否符合可测量要求。我们测试了有效宽度、纵向坡度、横向坡度和路面状况是否可以使用视觉语言模型(VLM)从行人视角图像中可靠评估。我们首次将基于采样的保形预测(CP)应用于基于VLM的无障碍评估。我们在来自韩国首尔的514张人行道图像上评估了四种VLM,并使用实地测量的地面真值。保形校准对所有模型和属性均达到了标称的90%覆盖率,但校准区域的信息量有所不同。有效宽度产生了最具信息量的估计,最佳模型的平均区间半宽约为1.0米。由于每个模型都高估了宽度,非对称校准在覆盖率不变的情况下将区间缩短了最多33%。纵向坡度仅具有边际信息量,横向坡度区间过宽而无法分辨监管阈值,路面状况集合在四种模型中的三种中退化为全部五个等级(A-E)。来自原始采样离散度的未校准区间在标称90%水平下仅覆盖了实地测量值的17-47%。在具有最自洽响应的图像中,这些区间在多达96%的情况下未包含实地测量值。因此,响应自洽性并不能证明准确性,采样离散度在针对实地测量地面真值进行校准之前不能解释为不确定性。没有定量属性达到通用合规评估所需的精度,但CP仅从校准数据就能识别哪些属性可以支持对远离阈值的路段进行筛查。我们发布了带注释的行人视角图像及其对应的实地测量属性值。
英文摘要:
An important component of urban accessibility, particularly for wheelchair users and people with reduced mobility, is sidewalk compliance with measurable requirements. We test whether effective width, longitudinal slope, cross slope, and pavement condition can be assessed reliably from pedestrian-level imagery using vision-language models (VLMs). We present the first application of sampling-based conformal prediction (CP) for VLM-based accessibility assessment. We evaluate four VLMs on 514 sidewalk images from Seoul, South Korea, with field-measured ground truth. Conformal calibration attains the nominal 90% coverage for all models and attributes, but the calibrated regions differ in informativeness. Effective width yields the most informative estimates, with a mean interval half-width of about 1.0 m for the best model. Since every model overestimates width, asymmetric calibration shortens the intervals by up to 33% at unchanged coverage. Longitudinal slope is marginally informative, cross-slope intervals are too wide to resolve regulatory thresholds, and pavement-condition sets degenerate to all five grades (A-E) for three of the four models. Uncalibrated intervals from raw sampling dispersion cover only 17-47% of field-measured values at a nominal 90% level. Among the images with the most self-consistent responses, these intervals miss the field-measured value in up to 96% of cases. Response self-consistency is therefore not evidence of accuracy, and sampling dispersion cannot be interpreted as uncertainty until it has been calibrated against field-measured ground truth. No quantitative attribute reaches the precision required for general compliance assessment, but CP identifies from calibration data alone which attributes can support screening of segments far from the thresholds. We release the annotated pedestrian-level images and their corresponding field-measured attribute values.