arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

超越像素相似性:基于任务感知的GAN合成声纳数据评估用于机器人感知

Beyond Pixel Similarity: Task-Aware Evaluation of GAN-Based Synthetic Sonar Data for Robotic Perception

Hannan Ejaz Keen, Muhammad Moazam Fraz, Karsten Berns

arXiv 2609.18100首次发表:更新:

发表机构

XITASO GmbH IT & Software Solutions; National University of Sciences and Technology (NUST); Technical University of Kaiserslautern(XITASO信息技术与软件解决方案有限公司; 国立科技大学; 凯泽斯劳滕工业大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本研究针对GAN合成声纳数据,发现像素级图像保真度指标(SSIM、PSNR、MSE)不能充分反映下游目标检测性能,提出并验证了任务感知评估的必要性,表明PatchGAN配置虽像素相似性较低但检测效果更优。

AI 中文摘要

合成数据可以降低机器人感知训练数据的采集和标注成本,但生成保留与下游感知相关特征的传感器观测仍然具有挑战性,尤其是在声纳图像方面。在本工作中,我们研究了传统图像保真度指标是否能充分反映GAN生成的合成声纳数据的下游感知性能。我们采用了一个Pix2Pix条件生成对抗网络,具有四种不同感受野的判别器配置:PixelGAN、PatchGAN-16、PatchGAN-70和ImageGAN。模型使用来自两个数据集的声纳图像进行训练,并使用传统图像保真度指标进行评估,包括结构相似性指数(SSIM)、峰值信噪比(PSNR)和均方误差(MSE)。为了用任务导向的评估补充这些像素级度量,我们仅使用真实声纳图像训练了YOLOX-S、YOLOX-L和Faster R-CNN检测器,随后在所有判别器配置下使用相同的测试样本和标注对GAN生成的图像进行评估。结果揭示了图像保真度与下游目标检测性能之间的差异:在SSIM、PSNR和MSE上取得最佳效果的配置并不总是产生最佳的检测性能。特别是,PatchGAN配置尽管没有达到最高的像素级相似性得分,却取得了强劲的下游检测结果。这些发现表明,对于所考虑的数据集和模型,仅靠像素级图像保真度指标可能无法始终捕捉合成声纳观测的任务相关真实性,并促使在用于机器人感知的合成传感器数据评估中采用任务感知的方法。

英文摘要

Synthetic data can reduce the cost of collecting and annotating training data for robotic perception, but generating sensor observations that preserve the characteristics relevant to downstream perception remains challenging, particularly for sonar imagery. In this work, we investigate whether conventional image-fidelity metrics adequately reflect the downstream perception performance of GAN-generated synthetic sonar data. We employ a Pix2Pix conditional generative adversarial network with four discriminator configurations characterized by different receptive fields: PixelGAN, PatchGAN-16, PatchGAN-70, and ImageGAN. The models are trained using sonar imagery from two datasets and evaluated using conventional image-fidelity metrics, including Structural Similarity Index (SSIM), Peak Signal-to-Noise Ratio (PSNR), and Mean Squared Error (MSE). To complement these pixel-level measures with task-oriented evaluation, YOLOX-S, YOLOX-L, and Faster R-CNN detectors are trained exclusively on real sonar imagery and subsequently evaluated on the GAN-generated images using identical test samples and annotations across all discriminator configurations. The results reveal a discrepancy between image-fidelity and downstream object-detection performance: the configuration achieving the best SSIM, PSNR, and MSE does not consistently yield the best detection performance. In particular, PatchGAN configurations achieve strong downstream detection results despite not achieving the highest pixel-level similarity scores. These findings suggest, for the datasets and models considered, pixel-level image-fidelity metrics alone may not consistently capture the task-relevant realism of synthetic sonar observations and motivate the use of task-aware evaluation for synthetic sensor data intended for robotic perception.

CommentsAccepted at Sim2Real and Classical Control: From Rigorous Theory to Data-Driven Robotics - IROS Workshop 2026

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑