arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-04 至 2025-12-04 共收录 53 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 8 篇

2511.18751 2025-12-04 cs.CL 88%

Robust Multimodal Sentiment Analysis of Image-Text Pairs by Distribution-Based Feature Recovery and Fusion

基于分布的特征恢复与融合的鲁棒多模态图像-文本对情感分析

Daiqing Wu, Dongbao Yang, Yu Zhou, Can Ma

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) TMCC, College of Computer Science, Nankai University(TMCC,南开大学计算机学院)

专题命中 图文多模态 :multimodal(title,abstract);image-text(title,abstract);分类 cs.CL

AI总结 本文提出DRF方法,通过特征队列和分布估计,实现对图像-文本对中低质量和缺失模态的鲁棒情感分析。

Comments Accepted by ACM MM 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.18214 2025-12-04 cs.CV cs.AI cs.CL cs.LG 87%

VLSU: Mapping the Limits of Joint Multimodal Understanding for AI Safety

VLSU:联合多模态理解在AI安全中的极限映射

Shruti Palaskar, Leon Gatys, Mona Abdelrahman, Mar Jacobo, Larry Lindsey, Rutika Moharir, Gunnar Lund, Yang Xu, Navid Shiee, Jeffrey Bigham, Charles Maalouf, Joseph Yitan Cheng

机构 * Apple(苹果公司)

专题命中 图文多模态 :multimodal(title,abstract);image-text(abstract);multimodal foundation model(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 VLSU通过细粒度分类和组合分析,揭示了多模态安全评估中联合理解的缺陷,为改进AI安全研究提供关键测试平台。

Comments 10 pages, 5 figures, 4 tables, detailed appendix. Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03542 2025-12-04 cs.CV cs.AI 84%

V-ITI: Mitigating Hallucinations in Multimodal Large Language Models via Visual Inference-Time Intervention

V-ITI: 通过视觉推理时间干预缓解多模态大语言模型中的幻觉

Nan Sun, Zhenyu Zhang, Xixun Lin, Kun Wang, Yanmin Shang, Naibin Gu, Shuohuan Wang, Yu Sun, Hua Wu, Haifeng Wang, Yanan Cao

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Baidu Inc.(百度公司)

专题命中 图文多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 V-ITI通过视觉推理时间干预框架,有效缓解多模态大语言模型中的视觉相关幻觉问题,同时保持任务性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03463 2025-12-04 cs.CV cs.AI cs.CL 82%

Text-Printed Image: Bridging the Image-Text Modality Gap for Text-centric Training of Large Vision-Language Models

文本打印图像:为以文本为中心的大型视觉-语言模型训练弥合图像-文本模态差距

Shojiro Yamabe, Futa Waseda, Daiki Shiono, Tsubasa Takahashi

机构 * Turing Inc.(图灵公司) Institute of Science Tokyo(东京科学研究院) The University of Tokyo(东京大学) Tohoku University(东北大学)

专题命中 图文多模态 :image-text(title,abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 本研究提出文本打印图像(TPI)技术,通过生成合成图像弥合图像-文本模态差距,提升以文本为中心的大型视觉-语言模型训练效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.19094 2025-12-04 cs.CV cs.AI 81%

SATORI-R1: Incentivizing Multimodal Reasoning through Explicit Visual Anchoring

SATORI-R1: 通过显式视觉锚定激励多模态推理

Chuming Shen, Wei Wei, Xiaoye Qu, Yu Cheng

机构 * Huazhong University of Science and Technology(华中科技大学) The Chinese University of Hong Kong(香港中文大学)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 SATORI-R1通过显式视觉锚定提升多模态推理,采用三个可验证阶段和VQA-Verify数据集,在VQA任务中实现15.7%的准确率提升。

Comments 21 pages, 8 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03445 2025-12-04 cs.CV cs.AI 73%

Multi-Aspect Knowledge-Enhanced Medical Vision-Language Pretraining with Multi-Agent Data Generation

多方面知识增强的医学视觉-语言预训练与多代理数据生成

Xieji Li, Siyuan Yan, Yingsheng Liu, H. Peter Soyer, Monika Janda, Victoria Mar, Zongyuan Ge

机构 * Department of Data Science and AI, Faculty of Information Technology, Monash University(数据科学与人工智能系,信息科技学院,墨尔本大学) Victorian Melanoma Service, Alfred Health(维多利亚黑色素瘤服务,阿尔弗雷德健康) Frazer Institute, The University of Queensland, Dermatology Research Centre(弗雷泽研究所,昆士兰大学,皮肤科研究中心)

专题命中 图文多模态 :cross-modal(abstract);image-text(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出一种多代理数据生成与多方面知识增强的医学视觉-语言预训练框架,通过提升数据质量和细粒度对齐,实现零样本性能的突破。

Comments 10 pages. Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03087 2025-12-04 cs.MM cs.AI 73%

When Harmful Content Gets Camouflaged: Unveiling Perception Failure of LVLMs with CamHarmTI

当有害内容被伪装时:通过CamHarmTI揭示LVLMs的感知失败

Yanhui Li, Qi Zhou, Zhihong Xu, Huizhong Guo, Wenhai Wang, Dongxia Wang

机构 * Zhejiang University(浙江大学)

专题命中 图文多模态 :multimodal(abstract);image-text(abstract);分类 cs.AI、cs.MM

AI总结 本文提出CamHarmTI基准,揭示LVLMs在识别伪装有害内容时的感知不足,并通过微调提升模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00887 2025-12-04 cs.CV 57%

Multilingual Training-Free Remote Sensing Image Captioning

多语言免训练 遥感图像描述生成

Carlos Rebelo, Gil Rocha, João Daniel Silva, Bruno Martins

机构 * INESC-ID(葡萄牙里斯本INESC-ID研究所) Instituto Superior Técnico(葡萄牙里斯本技术大学) Faculdade de Engenharia da Universidade do Porto(葡萄牙波尔图大学工程学院)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出一种无需训练的多语言遥感图像描述生成方法,通过检索增强提示和图基重新排序策略,实现跨语言的高效描述生成。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2510.13747 2025-12-04 cs.CV 90%

InteractiveOmni: A Unified Omni-modal Model for Audio-Visual Multi-turn Dialogue

InteractiveOmni: 一种用于音频-视觉多轮对话的统一多模态模型

Wenwen Tong, Hewei Guo, Dongchuan Ran, Jiangnan Chen, Jiefan Lu, Kaibin Wang, Keqiang Li, Xiaoxu Zhu, Jiakui Li, Kehan Li, Xueheng Li, Lumin Li, Chenxu Guo, Jiasheng Zhou, Jiandong Chen, Xianye Wu, Jiahao Wang, Silei Wu, Lei Chen, Hanming Deng, Yuxuan Song, Dinghao Zhou, Guiping Zhong, Ken Zheng, Shiyin Kang, Lewei Lu

机构 * SenseTime Research(商汤科技研究院)

专题命中 音频语音多模态 :audio-visual(title,abstract);omni-modal(title,abstract);multi-modal(abstract);cross-modal(abstract)

AI总结 InteractiveOmni是一种统一的多模态模型,通过多阶段训练策略提升多轮对话能力,提供高效的音频-视觉交互体验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.17282 2025-12-04 cs.AI cs.SD 88%

ERF-BA-TFD+: A Multimodal Model for Audio-Visual Deepfake Detection

ERF-BA-TFD+: 一种用于音频视觉深度伪造检测的多模态模型

Xin Zhang, Jiaming Chu, Jian Zhao, Yuchu Jiang, Xu Yang, Lei Jin, Chi Zhang, Xuelong Li

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title,abstract);分类 cs.AI

AI总结 ERF-BA-TFD+通过结合增强接收场和音频视觉融合,提出了一种多模态深度伪造检测模型,在DDL-AV数据集上实现了最先进的检测性能。

Comments The paper is withdrawn after discovering a flaw in the theoretical derivation presented in Section Method. The incorrect step leads to conclusions that are not supported by the corrected derivation. We plan to reconstruct the argument and will release an updated version once the issue is fully resolved

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.22826 2025-12-04 cs.CV 87%

Some Modalities are More Equal Than Others: Decoding and Architecting Multimodal Integration in MLLMs

某些模态比其他模态更平等:在MLLMs中解码和架构多模态整合

Tianle Chen, Chaitanya Chakka, Arjun Reddy Akula, Xavier Thomas, Deepti Ghadiyaram

机构 * Boston University(波士顿大学) Google DeepMind(谷歌DeepMind)

专题命中 音频语音多模态 :multimodal(title,abstract);multi-modal(abstract);cross-modal(abstract);audio-visual(abstract)

AI总结 本文研究了多模态大语言模型对矛盾模态的鲁棒性,提出模态对齐调优策略以提升多模态推理能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2210.17139 2025-12-04 cs.LO cs.DM math.LO 50%

Nested Sequents for Intuitionistic Grammar Logics via Structural Refinement

通过结构细化方法为直觉语法学逻辑构建嵌套序列为

Tim S. Lyon

专题命中 音频语音多模态 :multi-modal(abstract)

AI总结 本文通过结构细化方法,为直觉语法学逻辑构建了无割嵌套序列表演系统,并证明了其保守性、不可判定性和可判定子类。

Comments This paper is currently under review. arXiv admin note: text overlap with arXiv:2107.01998

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2512.03918 2025-12-04 cs.CV 57%

UniMo: Unifying 2D Video and 3D Human Motion with an Autoregressive Framework

UniMo:基于自回归框架统一2D视频与3D人体运动

Youxin Pang, Yong Zhang, Ruizhi Shao, Xiang Deng, Feng Gao, Xu Xiaoming, Xiaoming Wei, Yebin Liu

机构 * Tsinghua University(清华大学) Meituan(美团)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 UniMo通过自回归框架统一2D视频与3D人体运动,实现同时生成与理解,提升多模态联合建模能力。

Comments https://carlyx.github.io/UniMo/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03837 2025-12-04 cs.CV 57%

Heatmap Pooling Network for Action Recognition from RGB Videos

用于RGB视频中动作识别的热图池化网络

Mengyuan Liu, Jinfu Liu, Yongkang Jiang, Bin He

机构 * State Key Laboratory of General Artificial Intelligence, Peking University, Shenzhen Graduate School(国家通用人工智能重点实验室,北京大学深圳研究生院) Imaging Department, DJI Technology Co., Ltd(大疆技术创新有限公司影像部) TongJi University(同济大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种用于视频中动作识别的热图池化网络,通过反馈池化模块提取稳健且简洁的人体特征,并结合多模态数据提升识别性能。

Comments Final Version of IEEE Transactions on Pattern Analysis and Machine Intelligence

Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence (2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03687 2025-12-04 cs.CV 57%

Active Visual Perception: Opportunities and Challenges

主动视觉感知:机遇与挑战

Yian Li, Xiaoyu Guo, Hao Zhang, Shuiwang Li, Xiaowei Dai

机构 * College of Computer Science and Engineering, Guilin University of Technology(桂林理工大学计算机科学与工程学院) New Engineering Industry College, Putian University(莆田大学新工程产业学院)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文探讨了主动视觉感知在复杂环境中的应用与挑战,分析了其在机器人、自动驾驶等领域的潜力及面临的实时处理与多模态整合问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03580 2025-12-04 cs.CV cs.CR 57%

Dynamic Optical Test for Bot Identification (DOT-BI): A simple check to identify bots in surveys and online processes

动态光学测试用于机器人识别(DOT-BI):一种简单的检查以在调查和在线过程中识别机器人

Malte Bleeker, Mauro Gotsch

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 DOT-BI通过利用人类对运动的感知来识别自动化系统,通过隐藏数字的动态特性区分人类与机器人,实验表明其在调查和在线过程中具有高识别率和良好的用户体验。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.05661 2025-12-04 cs.CV 57%

Language-Driven Object-Oriented Two-Stage Method for Scene Graph Anticipation

基于语言的面向对象两阶段方法用于场景图预见

Xiaomeng Zhu, Changwei Wang, Haozhe Wang, Xinyu Liu, Fangzhen Lin

机构 * Department of Computer Science and Engineering, The Hong Kong University of Science and Technology(计算机科学与工程系,香港科学与技术大学) Key Laboratory of Computing Power Network and Information Security, Ministry of Education, Shandong Computer Science Center, Qilu University of Technology(计算能力网络与信息安全重点实验室,教育部,山东计算机科学中心,齐鲁大学) Academy of Interdisciplinary Studies, The Hong Kong University of Science and Technology(跨学科研究学院,香港科学与技术大学)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出基于语言的面向对象两阶段方法,通过时间一致性正则化预测对象集动态和关系轨迹,显著提升视频场景图预见性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 3 篇

2411.05826 2025-12-04 cs.CV cs.AI cs.LG 86%

From Pixels to Prose: Advancing Multi-Modal Language Models for Remote Sensing

从像素到 prose:推进遥感多模态语言模型

Xintian Sun, Benji Peng, Charles Zhang, Fei Jin, Qian Niu, Junyu Liu, Keyu Chen, Ming Li, Pohsun Feng, Ziqian Bi, Ming Liu, Xinyuan Song, Yichao Zhang

机构 * Simon Fraser University(西蒙弗雷泽大学) University of Minnesota - Twin Cities(明尼苏达大学双城分校) Kyoto University(京都大学) Georgia Institute of Technology(佐治亚理工学院) National Taiwan Normal University(台湾师范大学) Purdue University(普渡大学) Emory University(埃默里大学) The University of Texas at Dallas(德克萨斯大学达拉斯分校)

专题命中 跨模态检索 :multi-modal(title,abstract);MLLM(abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文探讨了多模态语言模型在遥感中的应用,分析了其技术基础、挑战及未来发展方向,强调了其在环境监测和灾害响应中的重要作用。

Comments 10 pages, 1 figure

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03514 2025-12-04 cs.IR cs.AI cs.CL cs.CV 85%

M3DR: Towards Universal Multilingual Multimodal Document Retrieval

M3DR:迈向通用多语言多模态文档检索

Adithya S Kolavi, Vyoman Jain

机构 * CognitiveLab(认知实验室)

专题命中 跨模态检索 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 M3DR提出了一种通用多语言多模态文档检索框架,通过对比训练实现跨语言和跨模态对齐,显著提升了多语言场景下的检索性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03276 2025-12-04 cs.LG 78%

Too Late to Recall: Explaining the Two-Hop Problem in Multimodal Knowledge Retrieval

太迟回忆:多模态知识检索中双跳问题的解释

Constantin Venhoff, Ashkan Khakzar, Sonia Joseph, Philip Torr, Neel Nanda

机构 * University of Oxford(牛津大学) McGill University(麦吉尔大学) Meta MATS

专题命中 跨模态检索 :multimodal(title,abstract)

AI总结 本文研究了多模态知识检索中双跳问题的影响,发现VLMs在处理视觉输入时过晚解决实体表示导致事实回忆性能下降,并提出通过修补和提示方法恢复性能。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 13 篇

2512.02895 2025-12-04 cs.CV 83%

MindGPT-4ov: An Enhanced MLLM via a Multi-Stage Post-Training Paradigm

MindGPT-4ov: 一种通过多阶段训练范式增强的多模态大语言模型

Wei Chen, Chaoqun Du, Feng Gu, Wei He, Qizhen Li, Zide Liu, Xuhao Pan, Chang Ren, Xudong Rao, Chenfeng Wang, Tao Wei, Chengjun Yu, Pengfei Yu, Yufei Zheng, Chunpeng Zhou, Pan Zhou, Xuhan Zhu

机构 * MindGPT-4o Team(MindGPT-4o 团队) Li Auto Inc.(利汽车公司)

专题命中 多模态生成 :MLLM(title,abstract);multimodal(abstract);分类 cs.CV

AI总结 MindGPT-4ov通过多阶段训练范式提升多模态大语言模型的性能与泛化能力,实现低成本高效率的训练与部署。

Comments 33 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2502.09854 2025-12-04 cs.CL cs.AI cs.LG 81%

Scaling Multimodal Search and Recommendation with Small Language Models via Upside-Down Reinforcement Learning

通过倒置强化学习扩展小语言模型以支持多模态搜索与推荐

Yu-Chen Lin, Sanat Sharma, Hari Manikandan, Jayant Kumar, Tracy Holloway King, Jing Zheng

机构 * Adobe(Adobe公司) Meta

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文通过倒置强化学习和合成数据蒸馏,利用小语言模型实现高效多模态搜索与推荐,显著降低推理延迟和内存开销。

Comments Accepted by ICDM 2025 MMSR

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03125 2025-12-04 cs.LG cs.AI 79%

Mitigating Intra- and Inter-modal Forgetting in Continual Learning of Unified Multimodal Models

缓解统一多模态模型持续学习中的模态内和模态间遗忘

Xiwen Wei, Mustafa Munir, Radu Marculescu

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出MoDE,通过解耦模态以缓解统一多模态模型中的模态内和模态间遗忘问题。

Comments NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03375 2025-12-04 cs.LG 78%

MAGE-ID: A Multimodal Generative Framework for Intrusion Detection Systems

MAGE-ID:一种多模态生成框架用于入侵检测系统

Mahdi Arab Loodaricheh, Mohammad Hossein Manshaei, Anita Raja

机构 * Department of Computer Science, Hunter College and The Graduate Center, City University of New York(计算机科学系,亨特学院和研究生中心,纽约市立大学)

专题命中 多模态生成 :multimodal(title,abstract)

AI总结 MAGE-ID通过多模态生成框架提升入侵检测系统的数据增强效果,实现更平衡和一致的多模态合成,显著提升检测性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03623 2025-12-04 cs.LG cs.AI physics.ao-ph 70%

The promising potential of vision language models for the generation of textual weather forecasts

视觉语言模型在生成文本天气预报中的巨大潜力

Edward C. C. Steele, Dinesh Mane, Emilio Monti, Luis Orus, Rebecca Chantrill-Cheyette, Matthew Couch, Kirstine I. Dale, Simon Eaton, Govindarajan Rangarajan, Amir Majlesi, Steven Ramsdale, Michael Sharpe, Craig Smith, Jonathan Smith, Rebecca Yates, Holly Ellis, Charles Ewen

机构 * Met Office(英国气象局) Amazon Web Services(亚马逊网络服务) University of East Anglia(东安格利亚大学)

专题命中 多模态生成 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI

AI总结 本文研究了视觉语言模型在直接生成文本天气预报中的潜力,通过视频编码的格网天气数据提升气象服务的生产效率与创新。

Comments 7 pages, 2 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15644 2025-12-04 cs.CV cs.AI cs.CR 62%

Can VLMs Detect and Localize Fine-Grained AI-Edited Images?

视觉语言模型能否检测并定位细粒度的人工智能编辑图像?

Zhen Sun, Ziyi Zhang, Zeren Luo, Zhiyuan Zhong, Zeyang Sha, Tianshuo Cong, Zheng Li, Shiwen Cui, Weiqiang Wang, Jiaheng Wei, Xinlei He, Qi Li, Qian Wang

机构 * Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州)) Ant Group(蚂蚁集团) Tsinghua University(清华大学) Shandong University(山东大学) Wuhan University(武汉大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出FragFake基准,首次系统研究VLMs在编辑图像分类和定位中的应用,发现微调模型在准确性上表现优异,同时探索了基于GRPO的RLVR训练方法。

Comments 14pages,19 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.04082 2025-12-04 cs.CV 57%

PosterCopilot: Toward Layout Reasoning and Controllable Editing for Professional Graphic Design

PosterCopilot: 向专业图形设计中的布局推理与可控编辑迈进

Jiazhe Wei, Ken Li, Tianyu Lao, Haofan Wang, Liang Wang, Caifeng Shan, Chenyang Si

机构 * PRLab, Nanjing University(南京大学PRLab) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 PosterCopilot通过三阶段训练策略和完整工作流程,实现专业图形设计中布局推理与可控编辑的提升。

Comments Project page: https://postercopilot.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.14206 2025-12-04 cs.LG cs.AI 57%

Challenges and Limitations of Generative AI in Synthesizing Wearable Sensor Data

生成式AI在合成可穿戴传感器数据中的挑战与局限性

Flavio Di Martino, Franca Delmastro

机构 * Institute for Informatics and Telematics of the Nationa Research Council (IIT-CNR)(意大利国家研究 council 信息与电信研究所)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.AI

AI总结 本文探讨了生成式AI在合成可穿戴传感器数据中的挑战与局限,评估了现有模型在多模态处理、长依赖性和条件生成方面的不足,并提出了未来研究方向以提升生成模型的实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03430 2025-12-04 cs.CV 57%

Label-Efficient Hyperspectral Image Classification via Spectral FiLM Modulation of Low-Level Pretrained Diffusion Features

通过低级预训练扩散特征的光谱FiLM调制实现标签高效的超光谱图像分类

Yuzhen Hu, Biplab Banerjee, Saurabh Prasad

机构 * University of Houston, Texas, USA(德克萨斯大学) Indian Institute of Technology Bombay, Mumbai, India(印度班加罗尔理工学院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出了一种基于预训练扩散模型的标签高效超光谱图像分类方法,通过光谱FiLM调制融合空间和光谱信息,提升分类性能。

Comments Accepted to the ICML 2025 TerraBytes Workshop (June 9, 2025)

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.03335 2025-12-04 cs.CV cs.LG 57%

Step-by-step Layered Design Generation

逐步分层设计生成

Faizan Farooq Khan, K J Joseph, Koustava Goswami, Mohamed Elhoseiny, Balaji Vasan Srinivasan

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出逐步分层设计生成方法,通过分层变化建模设计生成过程,引入新评估套件并验证其有效性。

Journal ref AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏