arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2608.30633cs.CVcs.LG

面向基于深度学习的灾后损毁评估的量子格拉斯曼-普吕克令牌混合方法

Quantum-Grassmann-Plucker Token Mixing for Deep Learning-Based Post-Disaster Damage Assessment

Kooroush Farahkhah, Umut Lagap, Taha Rezaei, Saman Ghaffarian

首次发表
浏览论文内容

中文总结 AI 辅助

本研究首次将格拉斯曼-普吕克令牌混合应用于计算机视觉,提出QGP与HQML-GP两种头模型,在xBD龙卷风数据集的见/未见事件测试中,QGP均取得最优分类性能,为灾后损毁评估提供了无注意力的替代方案。

中文摘要 AI 辅助

从卫星图像开展及时的灾后建筑损毁评估是关键的工程决策支持任务,但该任务仍受限于类别不平衡、中间损毁状态模糊以及跨事件迁移能力有限等问题。本研究首次将格拉斯曼-普吕克(Grassmann-Plucker, GP)令牌混合方法应用于计算机视觉领域,并为图像分类任务引入两种扩展方案:量子启发式格拉斯曼-普吕克(Quantum-inspired Grassmann-Plucker, QGP)头与混合量子机器学习格拉斯曼-普吕克(Hybrid Quantum Machine Learning Grassmann-Plucker, HQML-GP)头。GP头通过普吕克坐标编码令牌对形成的子空间,以此表征图像块令牌间的多尺度关系;QGP则用振幅衍生的概率特征丰富这些坐标,而HQML-GP将模拟量子电路生成的期望值纳入几何令牌表示。研究采用xBD龙卷风数据集的灾前灾后配对图像块,使用冻结的六通道Vision Transformer基础编码器(采用16×16像素图像块)进行处理。在相同的训练、 checkpoint 选择与评估协议下,将三种基于GP的头与多层感知机及Transformer基线模型进行对比:使用Joplin与Moore龙卷风样本开展模型开发与见事件测试,保留Tuscaloosa样本用于未见事件评估。QGP在两个测试集上的准确率与宏F1值均领先:见事件测试为83.46%与64.50%,未见事件测试为66.45%与52.70%;HQML-GP虽取得最高的验证宏F1值(65.63%),但在两个测试集上均未超越QGP,且每轮训练所需时间显著更长。这些结果表明,GP令牌混合方法可作为传统基于Transformer的令牌混合方法的一种有竞争力的无注意力替代方案,适用于配对卫星图像的损毁分类任务。

英文摘要

Timely post-disaster building damage assessment from satellite imagery is a critical engineering decision support task, yet it remains constrained by class imbalance, ambiguous intermediate damage states, and limited cross-event transferability. This study presents, to our knowledge, the first application of Grassmann-Plucker (GP) token mixing to computer vision and introduces two extensions for image classification: the Quantum-inspired Grassmann-Plucker (QGP) head and the Hybrid Quantum Machine Learning Grassmann-Plucker (HQML-GP) head. The GP head represents multiscale relationships among image patch tokens by encoding subspaces formed by token pairs with Plucker coordinates; QGP enriches these coordinates with amplitude-derived probability features, whereas HQML-GP incorporates expectation values generated by a simulated quantum circuit into the geometric token representation. Paired pre- and post-event image patches from the xBD tornado dataset were processed using a frozen six-channel Vision Transformer base encoder with 16 x 16-pixel patches. The three GP-based heads were compared with multilayer perceptron and Transformer baselines under identical training, checkpoint selection, and evaluation protocols. Joplin and Moore tornado samples were used for model development and seen-event testing, while Tuscaloosa was reserved for unseen-event evaluation. QGP led both test sets in accuracy and macro-F1: 83.46% and 64.50% for the seen events, and 66.45% and 52.70% for the unseen event. Although HQML-GP obtained the highest validation macro-F1 of 65.63%, it did not surpass QGP on either test set and required substantially more training time per epoch. These results establish GP token mixing as a competitive attention-free alternative to conventional Transformer-based token mixing for paired satellite image damage classification.

发表机构

  • Centre for Urban Resilience and Analytics (CURA), Georgia Institute of Technology (Georgia Tech)(佐治亚理工学院城市韧性与分析中心(CURA))
  • University of Exeter(埃克塞特大学)
  • University College London (UCL)(伦敦大学学院(UCL))

机构由 AI 辅助整理,请以论文原文为准。

↑