AI 中文总结
本文提出CM-MAE框架,通过软对比对齐损失等技术实现视觉-无线跨场景表征迁移,在DeepSense 6G数据集上提升了跨场景任务的准确率。
AI 中文摘要
同步的相机与无线测量通过不同物理通道观测同一场景,核心难点在于部署时学习的表征会因视角、流量、光照和传播几何的变化而失效。本文提出CM-MAE,一种用于跨场景表征迁移的自监督视觉-无线预训练框架。评估的真实数据模型仅使用RGB帧和DeepSense 6G中可用的64波束接收功率向量,预训练期间不使用光线追踪路径、校准深度或波束索引标签。其核心预训练项是软对比对齐损失,该损失未将同步图像-无线对作为唯一正样本对,而是从测量波束功率轮廓的相似性构建目标分布,因此具有相似方向响应的非相同样本不会被作为负样本强制分开。掩码联合解码器通过在模态丢弃下重构隐藏视觉块和无线角聚类提供互补局部目标。预训练后,差分率微调规则让新的融合头快速适应,同时编码器缓慢更新。在序列不相交的DeepSense 6G协议下,添加软对齐损失将匹配线性探针迁移准确率从24.88%提升至29.49%,温和融合微调在未见过的场景6-8上达到77.38%的Top-1准确率,可选的转导归一化适配达到78.69%。由于融合设置在推理时使用同期64波束功率向量,这些结果应被视为表征迁移诊断,而非主动波束预测或减少扫描的主张。
英文摘要
Synchronized camera and wireless measurements observe the same scene through different physical channels. The central difficulty is that a representation learned in one deployment can fail when viewpoint, traffic, illumination, and propagation geometry change. This paper presents CM-MAE, a self-supervised vision--wireless pretraining framework for cross-scenario representation transfer. The evaluated real-data model uses only RGB frames and the measured 64-beam received-power vector available in DeepSense 6G; it does not use ray-traced paths, calibrated depth, or beam-index labels during pretraining. Its central pretraining term is a \emph{soft contrastive alignment loss}. Instead of making the synchronized image--wireless pair the only positive pair, this loss builds a target distribution from similarities between measured beam-power profiles, so nonidentical samples with similar directional responses are not forced apart as false negatives. A masked joint decoder provides the complementary local objective by reconstructing hidden visual patches and wireless angular clusters under modality dropout. After pretraining, a differential-rate fine-tuning rule lets a new fusion head adapt quickly while the encoders move slowly. Under a sequence-disjoint DeepSense 6G protocol, adding the soft alignment loss improves a matched linear-probe transfer average from 24.88\% to 29.49\%. Mild fusion fine-tuning reaches 77.38\% Top-1 accuracy on unseen Scenarios 6--8, and optional transductive normalization adaptation reaches 78.69\%. Since the fusion setting uses the contemporaneous 64-beam power vector at inference, these results should be read as representation-transfer diagnostics, not as proactive beam-prediction or reduced-sweeping claims.