发表机构
Federal Highway Administration; Turner-Fairbank Highway Research Center(联邦公路管理局; 特纳-费尔班克公路研究中心)
机构由 AI 辅助整理,请以论文原文为准。AI 中文总结
本研究提出一种整合ResNet-50图像特征提取与GPT-2语言生成的VQA模型,用于自动化NDE图像分析,可提升检测效率、减少误差并增强现场可用性。
AI 中文摘要
本研究介绍一种专为无损检测(NDE)应用设计的视觉问答(VQA)模型。VQA模型允许检测人员与NDE图像进行交互,提出诸如“是否存在裂纹”或“缺陷位于何处”等针对性问题,并从模型获得精确答案。该系统利用深度学习与自然语言处理,整合了通过ResNet-50模型实现的图像特征提取,以及通过GPT-2实现的语言生成能力,以提供准确、信息丰富的反馈。通过实现直接的问答交互,此VQA模型显著提升了检测效率,减少了潜在误差,并增强了实际现场场景中的可用性。
英文摘要
An AI-based approach called ChatNDE Figure to Caption is introduced, which aims to automate the interpretation of NDE images using deep learning and natural language processing (NLP). A Vision-and-Language Pretraining (VLP) strategy is developed to help the model learn how to connect visual features with meaningful language. Basically, we built a large NDE image dataset, trained the model using annotated examples, and then evaluated how well it performed using BLEU scores to compare its output to expert written descriptions. So, the system combines a ResNet50 model to extract important features from the images and a GPT2 language model to turn those features into natural sounding text. Even though the accuracy of the model has been low the generated caption results have been solid so far, the captions were shorter but mentioned some important features of images what human experts would say, which shows the model is learning to pick up on key details. Also, a Visual Question Answering (VQA) model is used as part of the system. VQA models are designed to take an image and a question about that image (like Is there a crack? or Where is the defect located?) and generate a useful answer. By adding this layer, the platform will not just describe what it sees, it can also respond to specific questions, making it even more interactive and helpful for inspectors in the field. This whole approach is a big step toward speeding up NDE workflows, reducing human error, and making the technology more accessible.
Comments31, 20