发表机构
University of Illinois Urbana-Champaign(伊利诺伊大学厄巴纳-香槟分校)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
针对火灾后物体不可逆退化,提出TRACE基准和即插即用的特征恢复模块FRM,将退化特征映射为原始对齐表示,显著提升检测、检索、材料恢复、描述生成与功能推理性能。
AI 中文摘要
火灾后环境中的物体往往经历不可逆的物理转变,这些转变改变了它们的几何形状、材料状态和视觉外观。检测和识别这些残留物对于定位危险、重建事故前的内容以及清点损失至关重要。与标准图像损坏不同,这些退化影响的是物体本身的物理结构。为了研究这一场景,我们引入了TRACE,一个面向火灾后物体理解的变换感知基准。TRACE包含21.4K个基于真实图像的合成场景,以及成对的物体级原始到退化进展,涵盖189个类别中的499个物体身份。我们定义了五个针对定位和退化前理解的任务:退化物体检测、原始状态恢复与检索、原始材料恢复、原始描述生成以及功能推理。现有模型随严重程度的增加而急剧退化。从最轻到最严重的级别,RF-DETR的mAP相对下降了71%,而InternVL3.5的检索R@1从93.85降至28.11。为了解决这个问题,我们提出了特征恢复模块(FRM),一个即插即用的模块,将退化的编码器特征映射到原始对齐的表示,同时保持宿主冻结。仅使用成对特征监督进行训练,FRM提高了场景级检测、CLIP/SigLIP2特征恢复以及所有四个物体级VLM任务,且在更严重的退化下获得更大的提升。跨VLM宿主和严重程度级别,检索的相对增益平均为12.5%,材料恢复为20.1%,描述生成为13.2%,功能推理为12.4%。
英文摘要
Objects in post-fire environments often undergo irreversible physical transformations that change their geometry, material state, and visual appearance. Detecting and identifying these remnants is critical for locating hazards, reconstructing pre-incident contents, and inventorying losses. Unlike standard image corruptions, these degradations affect the physical structure of the object itself. To study this setting, we introduce TRACE, a transformation-aware benchmark for post-fire object understanding. TRACE contains 21.4K real-image-grounded synthetic scenes and paired object-level pristine-to-degraded progressions spanning 499 object identities across 189 categories. We define five tasks targeting localization and pre-degradation understanding: degraded-object detection, pristine-state recovery and retrieval, original material recovery, pristine description generation, and functional reasoning. Existing models degrade sharply with severity. From the least to the most severe level, RF-DETR mAP decreases by 71% relative, while InternVL3.5 retrieval R@1 falls from 93.85 to 28.11. To address this, we propose the Feature Recovery Module (FRM), a plug-and-play module that maps degraded encoder features to pristine-aligned representations while keeping the host frozen. Trained only with paired feature supervision, FRM improves scene-level detection, CLIP/SigLIP2 feature recovery, and all four object-level VLM tasks, with larger gains under more severe degradation. Across VLM hosts and severity levels, relative gains average 12.5% for retrieval, 20.1% for material recovery, 13.2% for description generation, and 12.4% for functional reasoning.
Comments28 pages, 11 figures, 9 tables