arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-02-11 至 2026-02-11 共收录 55 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2602.09080 2026-02-11 cs.LG cs.AI 79%

Looping Back to Move Forward: Recursive Transformers for Efficient and Flexible Large Multimodal Models

循环回溯以前进:递归变换器用于高效灵活的大型多模态模型

Ruihan Xu, Yuting Gao, Lan Wang, Jianing Li, Weihao Chen, Qingpei Guo, Ming Yang, Shiliang Zhang

机构 * State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,计算机学院,北京大学) AntGroup(蚂蚁集团)

专题命中 图文多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 RecursiveVLM通过递归细化机制提升多模态模型效率,实现参数复用和性能提升。

Comments This is a primary contribution in the Recursive Vision-Language Models

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.06218 2026-02-11 cs.CV cs.LG 79%

Cross-Modal Redundancy and the Geometry of Vision-Language Embeddings

跨模态冗余与视觉-语言嵌入的几何学

Grégoire Dhimoïla, Thomas Fel, Victor Boutin, Agustin Picard

机构 * Brown University(布朗大学) ENS Paris Saclay(巴黎萨克雷大学) IRT Saint Exupéry(IRT圣埃克苏佩里) Kempner Institute, Harvard University(哈佛大学凯姆纳研究所)

专题命中 图文多模态 :cross-modal(title,abstract);分类 cs.CV

AI总结 本文通过等能假设和对齐稀疏自编码器,揭示了视觉-语言模型中跨模态对齐的几何结构,发现稀疏双模态原子承载了跨模态对齐信号,单模态原子解释了模态差距,去除单模态原子可消除差距而不影响性能。

Comments Published as a conference paper at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.16596 2026-02-11 cs.CV cs.AI 62%

SHIELD: Suppressing Hallucinations In LVLM Encoders via Bias and Vulnerability Defense

SHIELD:通过偏差和脆弱性防御抑制LVLM编码器中的幻觉

Yiyang Huang, Liang Shi, Yitian Zhang, Yi Xu, Yun Fu

机构 * Department of Electrical and Computer Engineering, Northeastern University(电气与计算机工程系,东北大学) Khoury College of Computer Science, Northeastern University(计算机科学学院,东北大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 SHIELD通过减少统计偏差、对抗固有偏差和解决脆弱性,有效抑制LVLM编码器中的对象幻觉。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.08688 2026-02-11 cs.AI cs.CV 62%

STELAR-VISION: Self-Topology-Aware Efficient Learning for Aligned Reasoning in Vision

STELAR-VISION:基于拓扑意识的高效学习以实现对齐推理

Chen Li, Han Zhang, Zhantao Yang, Fangyi Chen, Zihan Wang, Anudeepsekhar Bolimera, Marios Savvides

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 STELAR-VISION通过拓扑意识训练框架提升视觉-语言模型的推理准确性和效率,实现更高效的多模态任务处理。

Comments This paper has been accepted at AAAI 2026. This is the author's extended version. The final version will appear in the official proceedings

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10913 2026-02-11 cs.CV cs.CL 62%

Know "No" Better: A Data-Driven Approach for Enhancing Negation Awareness in CLIP

了解“不”:一种数据驱动的方法用于增强CLIP中的否定意识

Junsung Park, Jungbeom Lee, Jongyoon Song, Sangwon Yu, Dahuin Jung, Sungroh Yoon

机构 * Department of Electrical and Computer Engineering, Seoul National University(首尔国立大学电子与计算机工程系) Amazon(亚马逊) Samsung Research(三星研究院) School of Computer Science and Engineering, Soongsil University(顺天大学计算机科学与工程学院) IPAI, AIIS, ASRI, INMC, and ISRC, Seoul National University(首尔国立大学IPAI、AIIS、ASRI、INMC和ISRC)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文提出NegationCLIP,通过生成包含否定的数据增强CLIP的否定意识,同时提出NegRefCOCOg基准用于评估多模态模型的否定理解能力。

Comments Accepted to ICCV 2025

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2025, pp. 2825-2835

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09443 2026-02-11 cs.AI 57%

P1-VL: Bridging Visual Perception and Scientific Reasoning in Physics Olympiads

P1-VL: 联结视觉感知与物理竞赛中的科学推理

Yun Luo, Futing Wang, Qianjia Cheng, Fangchen Yu, Haodi Lei, Jianhao Yan, Chenxi Li, Jiacheng Chen, Yufeng Zhao, Haiyuan Wan, Yuchen Zhang, Shenghe Zheng, Junchi Yao, Qingyang Zhang, Haonan He, Wenxuan Zeng, Li Sheng, Chengxing Xie, Yuxin Zuo, Yizhuo Li, Yulun Wu, Rui Huang, Dongzhan Zhou, Kai Chen, Yu Qiao, Lei Bai, Yu Cheng, Ning Ding, Bowen Zhou, Peng Ye, Ganqu Cui

机构 * Shanghai AI Laboratory(上海人工智能实验室)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

AI总结 P1-VL通过融合课程强化学习和代理增强技术,成为首个在物理竞赛中获得12枚金牌的开源视觉-语言模型,展示了卓越的科学推理能力和泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 4 篇

2510.21797 2026-02-11 cs.LG cs.AI cs.SD eess.AS 87%

Quantifying Multimodal Imbalance: A GMM-Guided Adaptive Loss for Audio-Visual Learning

量化多模态不平衡:一种基于GMM的自适应损失用于音频-视觉学习

Zhaocheng Liu, Zhiwen Yu, Xiaoqing Liu

机构 * South China University of Technology(南方科技大学)

专题命中 音频语音多模态 :multimodal(title,abstract);audio-visual(title);分类 cs.AI、eess.AS

AI总结 本文提出基于GMM的自适应损失,用于量化和缓解多模态学习中的样本级不平衡问题,通过双阶段框架提升音频-视觉学习性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09637 2026-02-11 cs.CV cs.MM 81%

Towards Training-free Multimodal Hate Localisation with Large Language Models

无需训练的多模态仇恨定位与大型语言模型

Yueming Sun, Long Yang, Jianbo Jiao, Zeyu Fu

机构 * Hybrid Intelligence Lab, University of Durham(杜伦大学混合智能实验室) Multimodal Intelligence Lab, University of Exeter(埃克塞特大学多模态智能实验室) The MIx Group University of Birmingham(伯明翰大学MIx集团)

专题命中 音频语音多模态 :multimodal(title);cross-modal(abstract);分类 cs.CV、cs.MM

AI总结 提出无需训练的多模态仇恨视频定位框架LELA,通过多阶段提示方案和跨模态推理机制实现高精度定位。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09121 2026-02-11 cs.AI 79%

Uncertainty-Aware Multimodal Emotion Recognition through Dirichlet Parameterization

通过狄利克雷参数化实现的不确定性感知多模态情绪识别

Rémi Grzeczkowicz, Eric Soriano, Ali Janati, Miyu Zhang, Gerard Comas-Quiles, Victor Carballo Araruna, Aneesh Jonelagadda

机构 * Kaliber Labs(Kaliber实验室)

专题命中 音频语音多模态 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出一种轻量级且隐私保护的多模态情绪识别框架,通过狄利克雷参数化方法协调不同模态的不确定性,实现高效且鲁棒的情绪识别。

Comments 8 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08794 2026-02-11 cs.CV cs.SD 77%

MOVA: Towards Scalable and Synchronized Video-Audio Generation

MOVA:迈向可扩展且同步的视频-音频生成

OpenMOSS Team, Donghua Yu, Mingshu Chen, Qi Chen, Qi Luo, Qianyi Wu, Qinyuan Cheng, Ruixiao Li, Tianyi Liang, Wenbo Zhang, Wenming Tu, Xiangyu Peng, Yang Gao, Yanru Huo, Ying Zhu, Yinze Luo, Yiyang Zhang, Yuerong Song, Zhe Xu, Zhiyu Zhang, Chenchen Yang, Cheng Chang, Chushu Zhou, Hanfu Chen, Hongnan Ma, Jiaxi Li, Jingqi Tong, Junxi Liu, Ke Chen, Shimin Li, Shiqi Jiang, Songlin Wang, Wei Jiang, Zhaoye Fei, Zhiyuan Ning, Chunguo Li, Chenhui Li, Ziwei He, Zengfeng Huang, Xie Chen, Xipeng Qiu

专题命中 音频语音多模态 :multimodal(abstract);image-text(abstract);audio-visual(abstract);分类 cs.CV

AI总结 MOVA通过混合专家架构生成高质量同步音频视频内容,支持图像-文本到视频-音频的生成任务,推动开放研究与创作者社区发展。

Comments Technical report for MOVA (open-source video-audio generation model). 38 pages, 10 figures, 22 tables. Project page: https://mosi.cn/models/mova Code: https://github.com/OpenMOSS/MOVA Models: https://huggingface.co/collections/OpenMOSS-Team/mova. Qinyuan Cheng and Tianyi Liang are project leader. Xie Chen and Xipeng Qiu are corresponding authors

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 5 篇

2506.18862 2026-02-11 cs.CV cs.AI 81%

TAMMs: Change Understanding and Forecasting in Satellite Image Time Series with Temporal-Aware Multimodal Models

TAMMs: 卫星图像时间序列中基于时序感知的多模态模型用于变化理解和预测

Zhongbin Guo, Yuhao Wang, Ping Jian, Chengzhi Li, Xinyue Chen, Zhen Yang, Ertai E

机构 * School of Computer Science & Technology, Beijing Institute of Technology(计算机科学与技术学院,北京理工大学) School of Computing, National University of Singapore(计算学院,新加坡国立大学)

专题命中 视频多模态 :multimodal(title);MLLM(abstract);分类 cs.CV、cs.AI

AI总结 TAMMs通过时序感知多模态模型统一执行卫星图像时间序列中的变化理解和未来预测,提升长程时间动态建模能力。

Comments Published as a conference paper at The Fourteenth International Conference on Learning Representations (ICLR 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09638 2026-02-11 cs.CV 79%

VideoAfford: Grounding 3D Affordance from Human-Object-Interaction Videos via Multimodal Large Language Model

VideoAfford: 通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding

Hanqing Wang, Mingyu Liu, Xiaoyu Chen, Chengwei MA, Yiming Zhong, Wenti Yin, Yuhao Liu, Zhiqing Cui, Jiahao Yuan, Lu Dai, Zhiyuan Ma, Hui Xiong

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 VideoAfford通过多模态大语言模型实现人类-物体交互视频中的3D affordance grounding,结合动态交互先验和空间感知损失函数,提升机器人操作的可操作区域识别能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.06303 2026-02-11 cs.LG 78%

Multimodal Graph Neural Networks for Prognostic Modeling of Brain Network Reorganization

多模态图神经网络用于脑网络重组的预后建模

Preksha Girish, Rachana Mysore, Kiran K. N., Hiranmayee R., Shipra Prashanth, Shrey Kumar

专题命中 视频多模态 :multimodal(title,abstract)

AI总结 本文提出多模态图神经网络用于脑网络重组的预后建模,通过整合多种影像数据,生成可解释的生物标志物以预测认知下降风险。

Comments Fundamental methodological error invalidating results

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09618 2026-02-11 cs.SI 50%

UniShare: A Unified Framework for Joint Video and Receiver Recommendation in Social Sharing

UniShare: 一种用于社交分享中视频和接收者推荐的统一框架

Caimeng Wang, Li Chong, Dongxu Liu, Xu Min, Jianhui Bu

专题命中 视频多模态 :multi-modal(abstract)

AI总结 UniShare通过统一框架联合预测视频和接收者推荐,利用增强的表示学习和联合训练范式提升分享效果,实现在快手平台上的显著性能提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20906 2026-02-11 cs.LG 50%

TwinWeaver: An LLM-Based Foundation Model Framework for Pan-Cancer Digital Twins

TwinWeaver: 一种基于大语言模型的跨癌症数字双胞胎基础模型框架

Nikita Makarov, Maria Bordukova, Lena Voith von Voithenberg, Estrella Pivel-Villanueva, Sabrina Mielke, Jonathan Wickes, Hanchen Wang, Mingyu Derek Ma, Keunwoo Choi, Kyunghyun Cho, Stephen Ra, Raul Rodriguez-Esteban, Fabian Schmich, Michael Menden

机构 * Computational Sciences Center of Excellence, Roche, Penzberg, Germany(罗氏计算科学卓越中心) Computational Health Center, Helmholtz Munich, Munich, Germany(海德堡慕尼黑计算健康中心) Department of Biology, Ludwig Maximilian University of Munich, Munich, Germany(慕尼黑路易斯·马克西米利安大学生物学系) Early Development Oncology, Roche Innovation Center Zurich, Roche, Schlieren, Switzerland(罗氏苏黎世创新中心早期肿瘤学) Computational Sciences Center of Excellence, Genentech, New York City, USA(基因泰克计算科学卓越中心) Computational Sciences Center of Excellence, Genentech, South San Francisco, USA(基因泰克计算科学卓越中心) Center for Data Science, New York University, New York City, USA(纽约大学数据科学中心) Department of Computer Science, Stanford University, Stanford, CA, USA(斯坦福大学计算机科学系) Computational Sciences Center of Excellence, Roche, Basel, Switzerland(罗氏巴塞尔计算科学卓越中心) Department of Biochemistry(生物化学系) Biotechnology Institute, The University of Melbourne, Melbourne, Australia(墨尔本大学生物技术研究所)

专题命中 视频多模态 :multi-modal(abstract)

AI总结 TwinWeaver通过基于大语言模型的框架,构建跨癌症数字双胞胎,提升临床事件预测与风险分层精度。

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 4 篇

2602.10023 2026-02-11 cs.CL 79%

MEVER: Multi-Modal and Explainable Claim Verification with Graph-based Evidence Retrieval

MEVER:基于图的证据检索的多模态和可解释性声明验证

Delvin Ce Zhang, Suhan Cui, Zhelin Chu, Xianren Zhang, Dongwon Lee

机构 * University of Sheffield(谢菲尔德大学) University of Science and Technology Beijing(北京科技大学) University of California San Diego(加州大学圣地亚哥分校) The Pennsylvania State University(宾夕法尼亚州立大学)

专题命中 跨模态检索 :multi-modal(title,abstract);分类 cs.CL

AI总结 MEVER通过多模态图检索和解释生成,实现了准确且可解释的声明验证,同时创建了AI领域的科学数据集AIChartClaim。

Comments Accepted to EACL-26

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09839 2026-02-11 cs.CV 79%

ARK: A Dual-Axis Multimodal Retrieval Benchmark along Reasoning and Knowledge

ARK:一个双轴多模态检索基准,沿推理与知识

Yijie Lin, Guofeng Ding, Haochen Zhou, Haobin Li, Mouxing Yang, Xi Peng

机构 * College of Computer Science, Sichuan University, Chengdu, China(四川大学计算机学院) National Key Laboratory of Fundamental Algorithms and Models for Engineering Simulation, Sichuan University, China(四川省工程仿真基础算法与模型国家重点实验室)

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.CV

AI总结 ARK基准通过双轴视角评估多模态检索,揭示知识密集型与推理密集型检索间的差距,发现细粒度视觉推理是主要瓶颈。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09401 2026-02-11 cs.IR 50%

SARM: LLM-Augmented Semantic Anchor for End-to-End Live-Streaming Ranking

SARM: 基于大语言模型的语义锚点用于端到端实时直播排序

Ruochen Yang, Yueyang Liu, Zijie Zhuang, Changxin Lao, Yuhui Zhang, Jiangxia Cao, Jia Xu, Xiang Chen, Haoke Xiao, Xiangyu Wu, Xiaoyou Zhou, Xiao Lv, Shuang Yang, Tingwen Liu, Zhaojie Liu, Han Li, Kun Gai

专题命中 跨模态检索 :multimodal(abstract)

AI总结 SARM通过整合自然语言语义锚点到端到端排序优化中,提升实时直播内容排序的精度和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09132 2026-02-11 cs.DB cs.ET cs.MA 50%

SciDataCopilot: An Agentic Data Preparation Framework for AGI-driven Scientific Discovery

SciDataCopilot: 一种面向AGI驱动科学发现的代理数据准备框架

Jiyong Rao, Yicheng Qiu, Jiahui Zhang, Juntao Deng, Shangquan Sun, Fenghua Ling, Hao Chen, Nanqing Dong, Zhangyang Gao, Siqi Sun, Yuqiang Li, Dongzhan Zhou, Guangyu Wang, Lijun Wu, Conghui He, Xuhong Wang, Jing Shao, Xiang Liu, Yu Zhu, Mianxin Liu, Qihao Zheng, Yinghui Zhang, Jiamin Wu, Xiaosong Wang, Shixiang Tang, Wenlong Zhang, Bo Zhang, Wanli Ouyang, Runkai Zhao, Chunfeng Song, Lei Bai, Chi Zhang

专题命中 跨模态检索 :multi-modal(abstract)

AI总结 SciDataCopilot通过自主代理框架实现端到端的数据准备,提升科学发现效率与一致性,推动实验驱动的科学通用智能发展。

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 9 篇

2509.22761 2026-02-11 cs.CV cs.AI 84%

MILR: Improving Multimodal Image Generation via Test-Time Latent Reasoning

MILR:通过测试时潜在推理改进多模态图像生成

Yapeng Mi, Yanpeng Zhao, Hengli Li, Chenxi Li, Huimin Wu, Xiaojian Ma, Song-Chun Zhu, Ying Nian Wu, Qing Li

机构 * University of Science and Technology of China(中国科学技术大学) State Key Laboratory of General Artificial Intelligence, BIGAI(通用人工智能国家重点实验室) Peking University(北京大学) Tsinghua University(清华大学) University of California, Los Angeles(加州大学洛杉矶分校)

专题命中 多模态生成 :multimodal(title,abstract);cross-modal(abstract);分类 cs.CV、cs.AI

AI总结 MILR通过测试时潜在推理提升多模态图像生成性能,实现跨模态推理和统一潜在空间优化。

Comments 21 pages,14 figures,9 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09528 2026-02-11 cs.CV 79%

SchröMind: Mitigating Hallucinations in Multimodal Large Language Models via Solving the Schrödinger Bridge Problem

SchröMind: 通过求解薛定谔桥问题减轻多模态大语言模型中的幻觉

Ziqiang Shi, Rujie Liu, Shanshan Yu, Satoshi Munakata, Koichi Shirahata

机构 * Fujitsu Research \& Development Center Co.,LTD., Beijing, China Fujitsu Limited, Tokyo, Japan

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 SchröMind通过求解薛定谔桥问题,有效减轻多模态大语言模型中的幻觉问题,提升模型在医疗等高风险领域的应用能力。

Comments ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.22274 2026-02-11 cs.CV cs.CL 76%

Common Objects Out of Context (COOCo): Investigating Multimodal Context and Semantic Scene Violations in Referential Communication

常见物体脱离上下文(COOCo):研究多模态上下文和语义场景违规在指代通信中的作用

Filippo Merlo, Ece Takmaz, Wenkai Chen, Albert Gatt

机构 * Utrecht University(乌特雷赫大学) University of Trento(特伦托大学)

专题命中 多模态生成 :multimodal(title);分类 cs.CV、cs.CL

AI总结 COOCo研究了VLMs在指代生成中如何利用场景上下文,发现模型根据语义相关性和噪声水平动态平衡局部与上下文信息。

Comments Accepted to TACL (pre-MIT Press publication version)

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.08682 2026-02-11 cs.CV 70%

ALIVE: Animate Your World with Lifelike Audio-Video Generation

ALIVE: 用逼真音频视频生成动画你的世界

Ying Guo, Qijun Gan, Yifu Zhang, Jinlai Liu, Yifei Hu, Pan Xie, Dongjun Qian, Yu Zhang, Ruiqi Li, Yuqi Zhang, Ruibiao Lu, Xiaofeng Mei, Bo Han, Xiang Yin, Bingyue Peng, Zehuan Yuan

机构 * Bytedance ALIVE Team(字节跳动ALIVE团队)

专题命中 多模态生成 :cross-modal(abstract);audio-visual(abstract);分类 cs.CV

AI总结 ALIVE通过结合预训练文本到视频模型与新的音频视频生成技术,实现了逼真的音频视频生成和动画,展示了优于开源和商业解决方案的性能。

Comments Technical report for ALIVE. Bytedance ALIVE Team. Homepage: https://foundationvision.github.io/Alive/

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.10113 2026-02-11 cs.CV 57%

ConsID-Gen: View-Consistent and Identity-Preserving Image-to-Video Generation

ConsID-Gen:视图一致且身份保持的图像到视频生成

Mingyang Wu, Ashirbad Mishra, Soumik Dey, Shuo Xing, Naveen Ravipati, Hansi Wu, Binbin Li, Zhengzhong Tu

机构 * Texas A&M University(德克萨斯大学) eBay Inc.(eBay公司)

专题命中 多模态生成 :cross-modal(abstract);分类 cs.CV

AI总结 ConsID-Gen通过视图辅助生成框架提升图像到视频生成的视图一致性和身份保持性。

Comments Project page: https://myangwu.github.io/ConsID-Gen

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09775 2026-02-11 cs.CV 57%

Where Do Images Come From? Analyzing Captions to Geographically Profile Datasets

图像来自哪里?通过分析描述词地理上剖析数据集

Abhipsa Basu, Yugam Bahl, Kirti Bhagat, Preethi Seshadri, R. Venkatesh Babu, Danish Pruthi

机构 * Indian Institute of Science, Bangalore(印度科学研究院,班加罗尔) TNSQ AI(TNSQ人工智能) University of California, Irvine(加州大学尔湾分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 研究通过分析图像描述词地理分布,发现训练数据严重偏向欧美国家,且代表性与GDP正相关,但生成图像的多样性不足。

Comments 41 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09432 2026-02-11 cs.CV 57%

SceneReVis: A Self-Reflective Vision-Grounded Framework for 3D Indoor Scene Synthesis via Multi-turn RL

SceneReVis: 一种基于多轮强化学习的视觉 grounding 框架,用于通过多轮强化学习进行 3D 室内场景合成

Yang Zhao, Shizhao Sun, Meisheng Zhang, Yingdong Shi, Xubo Yang, Jiang Bian

机构 * Shanghai Jiao Tong University, Shanghai, China(上海交通大学) Microsoft Research Asia, Beijing, China(微软亚洲研究院) Peking University, Beijing, China(北京大学) ShanghaiTech University, Shanghai, China(上海科技大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 SceneReVis 通过多轮强化学习和多模态反馈,实现高保真 3D 室内场景合成,提升空间规划能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02380 2026-02-11 cs.CV 57%

Unified Personalized Reward Model for Vision Generation

面向视觉生成的统一个性化奖励模型

Yibin Wang, Yuhang Zang, Feng Han, Jiazi Bu, Yujie Zhou, Cheng Jin, Jiaqi Wang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Jiaotong University(上海交通大学) Shanghai AI Lab(上海人工智能实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出UnifiedReward-Flex,一种面向视觉生成的统一个性化奖励模型,通过结合奖励建模与灵活上下文适应的推理,提升视觉生成的准确性与对人类偏好的对齐性。

Comments Website: https://codegoat24.github.io/UnifiedReward/flex

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.08701 2026-02-11 q-bio.QM cs.CV 57%

Automated Lesion Segmentation of Stroke MRI Using nnU-Net: A Comprehensive External Validation Across Acute and Chronic Lesions

利用nnU-Net实现脑卒中MRI病变自动分割:在急性与慢性病变上的全面外部验证

Tammar Truzman, Matthew A. Lambon Ralph, Ajay D. Halai

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文利用nnU-Net框架对脑卒中MRI病变进行自动分割,验证了模型在急性与慢性病变中的泛化能力,发现病变体积和图像质量是影响分割准确性的关键因素。

Comments 32 pages, 7 figures. Submitted to Brain. Code and trained models available

Journal ref 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 9 篇

2512.19379 2026-02-11 cs.LG cs.AI cs.MM 84%

OmniMER: Auxiliary-Enhanced LLM Adaptation for Indonesian Multimodal Emotion Recognition

OmniMER: 增辅增强的LLM适应用于印度尼西亚多模态情感识别

Xueming Yan, Boyan Xu, Yaochu Jin, Lixian Xiao, Wenlong Ye, Runyang Cai, Zeqi Zheng, Jingfa Liu, Aimin Yang, Yongduan Song

机构 * School of Information Science and Technology, Guangdong University of Foreign Studies(广东外语外贸大学信息科学与技术学院) School of Computer Science, Guangdong University of Technology(广东工业大学计算机学院) Faculty of Asian Languages and Cultures, Guangdong University of Foreign Studies(广东外语外贸大学亚洲语言文化学院) School of Engineering, Westlake University(西湖大学工程学院) School of Computer Science and Intelligence Education, Lingnan Normal University(岭南师范学院计算机科学与智能教育学院) School of Automation, Chongqing University(重庆大学自动化学院)

专题命中 多模态评测 :multimodal(title,abstract);cross-modal(abstract);分类 cs.AI、cs.MM

AI总结 OmniMER通过三种辅助任务提升印度尼西亚多模态情感识别性能,实现情感分类和识别的显著提升。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.09315 2026-02-11 cs.CV cs.AI 81%

A Deep Multi-Modal Method for Patient Wound Healing Assessment

一种用于患者伤口愈合评估的深度多模态方法

Subba Reddy Oota, Vijay Rowtula, Shahid Mohammed, Jeffrey Galitz, Minghsun Liu, Manish Gupta

机构 * Woundtech Innovative Healthcare Solutions(Woundtech创新医疗解决方案) Microsoft AI Research(微软人工智能研究院)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV、cs.AI

AI总结 本文提出了一种深度多模态方法,通过整合伤口变量和图像数据,预测患者住院风险,以提高伤口诊断效率和早期发现愈合问题。

Comments 4 pages, 2 figures

Journal ref Medical Imaging Meets NeurIPS Workshop, 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada

详情

展开后加载摘要…

URL PDF HTML 收藏