AI 中文总结
该研究针对城市洪水缺乏厘米级实时深度估计的问题,提出三种视觉-语言模型,通过机械可解释性分析确定关键交叉注意力层进行选择性微调,在合成与真实基准测试中均实现高精度洪水深度估计,大幅减少可训练参数。
AI 中文摘要
城市洪水对交通基础设施的威胁日益加剧,但目前尚无运行系统能提供街道路面级、厘米级分辨率的实时洪水深度估计。本文提出三种针对街景图像连续洪水深度估计微调的视觉-语言模型:FloodLlama-Dense(全微调的QLoRA基线模型)、FloodLlama-MI5与FloodLlama-MI6(由机械可解释性引导的稀疏变体,分别仅微调通过机械可解释性分析确定的前5层和6层因果相关交叉注意力层)。训练使用Unreal Engine 5生成的281万张合成图像中的约61万张子集,该数据集包含单车辆子集(深度增量5厘米)与多车辆子集(深度增量1厘米),涵盖7种车辆类型、4种天气条件,洪水深度范围为0至40厘米。FloodLlama-Dense的平均绝对误差(MAE)为0.40厘米、均方根误差(RMSE)为1.97厘米、5厘米精度(Acc@5cm)为97.59%。结合线性探测、logit lens、中心核对齐(CKA)与交叉注意力熵的机械可解释性分析揭示了两阶段适应模式:L13-L22层重构视觉表征,深度信息在L23层首次可线性解码。FloodLlama-MI5与FloodLlama-MI6利用该见解,仅微调8层交叉注意力层中的5或6层,可减少86%-88%的可训练参数(655万-786万,对比FloodLlama-Dense的5440万)且精度损失极小;在真实基准测试中,FloodLlama-MI6精度达98.62%,而已发布的STURM-FloodDepth基线模型精度为86.61%。
英文摘要
Urban flooding poses an escalating threat to transportation infrastructure, yet no operational system provides real-time, street-level flood-depth estimates at centimeter resolution. This paper presents three vision-language models fine-tuned for continuous flood-depth estimation from street-level imagery: FloodLlama-Dense, a fully fine-tuned QLoRA baseline, and FloodLlama-MI5 and FloodLlama-MI6, interpretability-guided sparse variants that fine-tune only the top five and six causally relevant cross-attention layers identified through mechanistic interpretability analysis, respectively. Training uses an approximately 610,000-image subset of a 2.81-million-image synthetic corpus generated in Unreal Engine 5. The dataset combines single-vehicle subsets with 5 cm depth increments and mixed-vehicle subsets with 1 cm depth increments, spanning seven vehicle types, four weather conditions, and flood depths from 0 to 40 cm. FloodLlama-Dense achieves an MAE of 0.40 cm, an RMSE of 1.97 cm, and an Acc@5cm of 97.59%. Mechanistic interpretability analysis combining linear probing, logit lens, centered kernel alignment (CKA), and cross-attention entropy reveals a two-stage adaptation pattern: layers L13-L22 restructure visual representations, while depth first becomes linearly decodable at layer L23. FloodLlama-MI5 and FloodLlama-MI6 leverage this insight by fine-tuning only five or six of the eight cross-attention layers, achieving an 86-88% reduction in trainable parameters (6.55-7.86 million versus 54.4 million) with minimal accuracy loss. On a real-world benchmark, FloodLlama-MI6 achieves 98.62% accuracy, compared with 86.61% for the published STURM-FloodDepth baseline.
CommentsThis Paper is accepted in International Conference on Machine Learning and Application (ICMLA) 2026