Multi-Modal Properties and Dynamics of the Gradient Echo Quantum Memory
专题命中 其他多模态 :multi-modal(title)
Comments 4 pages 3 figures
Journal ref Phys. Rev. Lett. 101, 203601 (2008)
AI 大模型
跨文本、图像、视频、音频等模态的大模型与学习方法。
专题命中 其他多模态 :multi-modal(title)
Comments 4 pages 3 figures
Journal ref Phys. Rev. Lett. 101, 203601 (2008)
FiLoRA: 专注与忽略LoRA用于可控的特征依赖
机构 * University of Melbourne, Melbourne, Australia ; Brain Science Institute, Korea Institute of Science ; Department of Computer Science ; Engineering, Korea University, Seoul, Republic of Korea ; Division of Bio-Medical Science \& Technology, University of Science ; Technology KIST School, Seoul, Republic of Korea
专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI
AI总结 FiLoRA通过指令条件门控实现对多模态模型内部特征依赖的可控调节,提升模型在虚假特征干预下的鲁棒性。
基于区域级组织推理与非平衡最优传输的淋巴细胞模拟修正
机构 * Duke University(杜克大学)
专题命中 其他多模态 :MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 针对细胞模拟导致的病理细胞分类歧义问题,提出Loki-OT方法,结合区域级组织推理与非平衡最优传输,在TCGA-BRCA队列上取得更优性能。
Comments 13 pages, 3 figures. Accepted to the MICCAI 2026 COMPAYL Workshop
Vernata:激光雷达点表示的自监督学习
机构 * Robotics and AI Institute(机器人与人工智能研究院) ; ETH Zurich(苏黎世联邦理工学院)
专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 本研究提出基于Sonata架构的Vernata自监督学习框架,通过三项扩展优化激光雷达点云表示,在多数据集上相比基线取得显著性能提升,模态减少场景下仍保持竞争力
Comments IROS 2026. Implementation: https://github.com/rai-opensource/vernata
视觉定位:一项综述
专题命中 其他多模态 :multimodal(abstract,abstract_cn);分类 cs.CV
AI总结 本综述梳理视觉定位的发展与背景,总结近年进展与新挑战,定义规范研究设置,介绍相关数据集与应用,提出未来方向,是该领域最全面的综述,适合不同阶段研究者。
Comments Accepted by TPAMI 2025. We keep tracing related works at https://github.com/linhuixiao/Awesome-Visual-Grounding, article publication page: https://ieeexplore.ieee.org/abstract/document/11235566
Journal ref IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 48, no. 3, pp. 2749-2771, March 2026
近视防控3.0:人工智能驱动的风险分层、主动监测和个性化干预
机构 * Shanghai Nile Intelligent Technology Co., Ltd.(上海尼罗智能科技有限公司) ; Beijing Tanyuan Academy of Intelligent Sensing(北京潭渊智能传感研究院)
专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI
AI总结 研究借助人工智能、数字传感等技术推动近视防控从被动变主动精准模式。通过多模态数据机器学习预测风险、可穿戴等监测及个性化干预形成闭环。评估各阶段证据,讨论相关挑战并给出未来方向。
ExACT: 基于示例驱动的校准精化用于遥感图像中免训练的视觉定位
机构 * Xidian University(西安电子科技大学)
专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 提出ExACT框架,通过一次性视觉提示机制弥合多模态大语言模型在遥感视觉定位中的模态差距,实现免训练的精确像素级定位。
Comments 11 pages, 8 figures, supplementary material included
改变模态:将遥感模型适应新卫星和传感器
机构 * University of British Columbia(不列颠哥伦比亚大学) ; Vector Institute(向量研究所) ; Carleton University(卡尔顿大学)
专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
AI总结 针对遥感模型在新卫星和传感器上的部署问题,提出DeluluNet架构,通过模态幻觉实现模态迁移、添加和子集三种场景下的模型适应,无需重新标注。
Comments 17 pages, 7 figures, 9 tables
Oracle剪枝真的是真正的Oracle吗?
机构 * Westlake University(西湖大学) ; Nankai University(南开大学) ; ENCODE Lab, Westlake University(西湖大学ENCODE实验室)
专题命中 其他多模态 :MLLM(abstract,abstract_cn);分类 cs.CV
AI总结 本文通过大规模实验(37K模型)发现,对于中等规模以上的深度学习模型,Oracle剪枝选择的权重在重训练后性能与重训练前几乎无关,质疑了Oracle剪枝作为剪枝方法基础的有效性。
Comments TMLR, Webpage: https://fscdc.github.io/Oracle-Pruning-Sanity-Check/
MLLMs 先正确后错误:追踪并纠正后层文本偏见
机构 * National University of Defense Technology(国防科技大学) ; Chinese University of Hong Kong, Shenzhen(香港中文大学(深圳)) ; Intelligent Game and Decision Lab(智能博弈与决策实验室)
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract_cn);分类 cs.CV
AI总结 发现多模态大语言模型在中间层形成正确视觉预测,但最终输出时被文本覆盖,通过检测预测方向变化(85%失败转向文本,89%成功转向视觉)提出无训练方法CALRD,在冲突基准上提升高达9.4%。
Comments Accepted at IJCAI 2026. 16 pages, 10 figures
MindZero:零标注的在线心智推理学习
机构 * University of Science and Technology of China(中国科学技术大学)
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.AI
AI总结 提出MindZero框架,通过自监督强化学习训练多模态大语言模型,实现高效鲁棒的在线心智推理,无需显式心智状态标注。
Comments ICML 2026. Website: https://scai.cs.jhu.edu/MindZero
基于推理的异常检测与定位:图像级监督
机构 * State Key Laboratory of Virtual Reality Technology and Systems, Beihang University(北京航空航天大学虚拟现实技术与系统国家重点实验室) ; Hangzhou Innovation Institute, Beihang University(北京航空航天大学杭州创新研究院)
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV
AI总结 本文提出基于图像级监督的异常检测与定位方法,利用大语言模型的推理能力实现像素级定位,无需额外组件或标注。
Comments Accepted to CVPR 2026
AERR-Nav:面向零样本物体导航的自适应探索-恢复-回忆策略
机构 * Hong Kong Polytechnic University(香港理工大学) ; Institute of automation, Chinese Academy of Sciences(中国科学院自动化研究所) ; Hong Kong University of Science and Technology (Guangzhou)(香港科技大学(广州))
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV
AI总结 本文提出AERR-Nav框架,通过自适应探索-恢复-回忆策略和自适应探索状态,解决零样本物体导航中探索与利用的平衡问题,在HM3D和MP3D基准测试中取得最佳性能。
多模态联邦学习用于磁共振成像图像分割
机构 * School of Artificial Intelligence(人工智能学院) ; School of Computer Science and Technology(计算机科学与技术学院) ; State Key Laboratory of Opto-Electronic Information Acquisition and Protection Technology(光电信息采集与防护技术国家重点实验室) ; Anhui Provincial Key Laboratory of Security Artificial Intelligence(安徽省安全人工智能重点实验室) ; Anhui Provincial Key Laboratory of Multimodal Cognitive Computation(安徽省多模态认知计算重点实验室)
专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文提出了一种多模态联邦学习框架,用于解决MRI图像分割中的模态异质性和数据异质性问题。
通过塑造密集且准确的2D语义预测来增强3D激光雷达分割
机构 * The University of Tokyo(东京大学)
专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
AI总结 本文提出MM2D3D模型,通过多模态引导滤波和动态跨伪监督提升2D预测质量,从而增强3D激光雷达分割的准确性。
HIPPO:通过混合模态偏好优化增强大语言模型的表格理解能力
机构 * School of Computer Science and Engineering, Northeastern University, Shenyang, China(东北大学计算机科学与工程学院) ; Department of Computer Science and Technology, Tsinghua University, Beijing, China(清华大学计算机科学与技术系) ; Huawei Technologies Co., Ltd(华为技术有限公司)
专题命中 其他多模态 :multi-modal(abstract);MLLM(abstract);分类 cs.CL
AI总结 HIPPO 通过混合模态偏好优化提升大语言模型的表格理解能力,实现 4% 的性能提升。
V-Loop:用于医学视觉问答中幻觉检测的视觉逻辑循环验证
机构 * Northwestern Polytechnical University(西北工业大学)
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV
AI总结 V-Loop通过双向推理和视觉逻辑循环验证,提升医学视觉问答中幻觉检测的准确性和效率。
Sketch-in-Latents: 在潜在空间中实现多模态统一推理
机构 * School of Computer Science and Technology, Huazhong University of Science and Technology(华中科技大学计算机科学与技术学院) ; Alibaba Cloud Computing(阿里巴巴云计算)
专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
AI总结 SkiLa通过在潜在空间中实现多模态统一推理,扩展MLLMs的自回归能力,生成连续视觉嵌入,提升视觉任务性能和多模态泛化能力。
Comments 14 pages, 11 figures
VGent: 通过模块化设计实现视觉 grounding 的解耦推理与预测
机构 * University of Illinois Chicago(伊利诺伊大学芝加哥分校) ; Adobe(Adobe公司)
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV
AI总结 VGent 通过模块化设计实现视觉 grounding 的解耦推理与预测,利用冻结 MLLM 和解码器交叉注意力机制,提升多目标识别性能。
Comments 8 pages
机构 * Rutgers University(罗格斯大学) ; Tsinghua University(清华大学) ; Michigan State University(密歇根州立大学) ; Florida State University(佛罗里达州立大学)
专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2025
专题命中 其他多模态 :multimodal(abstract);multimodal foundation model(abstract);分类 cs.AI
机构 * Sun Yat-sen University(中山大学) ; The Hong Kong University of Science and Technology (Guangzhou)(香港科学与技术大学(广州)) ; Peng Cheng Laboratory(鹏城实验室) ; Key Laboratory of Machine Intelligence and Advanced Computing, MOE(教育部机器智能与先进计算重点实验室)
专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
专题命中 其他多模态 :multi-modal(abstract);cross-modal(abstract);分类 cs.CV
Comments 12 pages, 3 figures
专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.CV
Comments Oral at ISMRM 2025
机构 * Soochow University(苏霍沃大学) ; Hong Kong University of Science and Technology(香港科技大学) ; Tencent(腾讯) ; Zhejiang Normal University(浙江师范大学)
专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.AI
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.CV
Comments Accept to CVPRW2025 (FGVC12)
专题命中 其他多模态 :multimodal(abstract);multi-modal(abstract);分类 cs.AI
专题命中 其他多模态 :multimodal(abstract);MLLM(abstract);分类 cs.AI
专题命中 其他多模态 :multimodal(abstract);cross-modal(abstract);分类 cs.CV
Comments Preprint
专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.CL、cs.AI
Journal ref EACL2023