arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.39964cs.AIcs.MAeess.SP

AIMS:用于仿真到现实多模态ISAC的智能体AI框架

AIMS: An Agentic AI Framework for Sim-to-Real Multi-Modal ISAC

发表机构香港科技大学
查看机构详情
  • The Hong Kong University of Science and Technology(香港科技大学)

机构由 AI 辅助整理,请以论文原文为准。

Yijie Bian, Kai Zhang, Wei Guo, Zixin Wang, Shenghui Song, Jun Zhang, Khaled B. Letaief

首次发表
浏览论文内容

中文总结 AI 辅助

提出AIMS智能体AI框架,通过双智能体架构和结构化知识实现仿真到现实多模态ISAC的部署配置生成与任务学习,在DeepSense 6G数据集上提升车辆检测和波束预测性能。

中文摘要 AI 辅助

多模态集成感知与通信(ISAC)使智能无线网络具备环境感知和可靠连接能力。数据驱动的多模态ISAC模型严重依赖带注释的真实世界数据来学习感知与无线观测之间的关系,从而限制了可扩展部署。尽管合成数据生成减轻了负担,但将现有仿真流程适配到目标部署需要一致的场景、感知、无线和学习配置,而这些耦合组件之间的不匹配会损害仿真到现实的迁移性。为应对这一挑战,我们提出了一种用于仿真到现实多模态ISAC的智能体人工智能(AI)框架,命名为AIMS。给定一个指定目标任务、部署条件和真实数据预算的自然语言部署请求,AIMS推导出特定于部署的仿真到现实配置,并协调其执行以生成特定于部署的任务模型。一种双智能体架构协调场景构建与任务学习。场景构建智能体从共享物理状态生成地理上接地、同步的感知和无线记录,而场景理解智能体配置与任务相关的模态和专家混合(MoE)学习,用于零样本推理或少样本适应。结构化领域知识指导依赖感知规划,而验证证据支持对受影响决策的反馈驱动修订。在真实世界DeepSense 6G数据集上的实验表明,与所考虑的仿真和融合基线相比,车辆检测和波束预测性能有所提升。一个单独的编排基准评估了跨多样部署请求的任务解释、依赖推理和反馈驱动重规划,显示结构化领域知识和验证反馈提高了规划正确性。

英文摘要

Multi-modal integrated sensing and communication (ISAC) enables environmental perception and reliable connectivity for intelligent wireless networks. Data-driven multi-modal ISAC models depend heavily on annotated real-world data to learn relationships across sensing and wireless observations, thereby constraining scalable deployment. Although synthetic data generation reduces the burden, adapting existing simulation pipelines to a target deployment requires consistent scene, sensing, wireless, and learning configurations, while mismatches among these coupled components impair sim-to-real transferability. To address the challenge, we propose an agentic artificial intelligence (AI) framework for sim-to-real multi-modal ISAC, named AIMS. Given a natural-language deployment request specifying the target task, deployment conditions, and real-data budget, AIMS derives a deployment-specific sim-to-real configuration and coordinates its execution to produce a deployment-specific task model. A two-agent architecture coordinates scene construction with task learning. A scene construction agent generates geographically grounded, synchronized sensing and wireless records from shared physical states, while a scene understanding agent configures task-relevant modalities and mixture-of-experts (MoE) learning for zero-shot inference or few-shot adaptation. Structured domain knowledge guides dependency-aware planning, while validation evidence supports feedback-driven revision of affected decisions. Experiments on the real-world DeepSense 6G dataset demonstrate improved vehicle detection and beam prediction over the considered simulation and fusion baselines. A separate orchestration benchmark evaluates task interpretation, dependency reasoning, and feedback-driven replanning across diverse deployment requests, showing improved plan correctness with structured domain knowledge and validation feedback.

↑