arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-12 至 2026-01-12 共收录 5 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 5 篇

2601.05572 2026-01-12 cs.CV 79%

Towards Generalized Multi-Image Editing for Unified Multimodal Models

面向统一多模态模型的通用多图像编辑

Pengcheng Xu, Peng Tang, Donghao Luo, Xiaobin Hu, Weichu Cui, Qingdong He, Zhennan Chen, Jiangning Zhang, Charles Ling, Boyu Wang

机构 * Western University(西方大学) Tencent YouTu Lab(腾讯优图实验室) Nanjing University(南京大学)

专题命中 多模态生成 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出了一种面向统一多模态模型的通用多图像编辑框架,通过可学习的潜在分离器和正弦索引编码提升多图像编辑任务中的视觉一致性和泛化能力。

Comments Project page: https://github.com/Pengchengpcx/MIE-UMM

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.05254 2026-01-12 cs.CV cs.AI cs.LG 62%

Explaining Low Perception Model Competency with High-Competency Counterfactuals

用高能力反事实解释低感知模型能力

Sara Pohland, Claire Tomlin

机构 * University of California, Berkeley(加州大学伯克利分校)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出五种生成高能力反事实图像的方法,用于解释模型预测不确定性的原因,并通过实验验证了其在生成语言解释中的有效性。

Journal ref Explainable Artificial Intelligence. xAI 2025. Communications in Computer and Information Science, vol 2580

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.05546 2026-01-12 cs.CV 57%

MoGen: A Unified Collaborative Framework for Controllable Multi-Object Image Generation

MoGen:一种用于可控多对象图像生成的统一协作框架

Yanfeng Li, Yue Sun, Keren Fu, Sio-Kei Im, Xiaoming Liu, Guangtao Zhai, Xiaohong Liu, Tao Tan

机构 * Faculty of Applied Sciences, Macao Polytechnic University(澳门理工学院) College of Computer Science, Sichuan University(四川大学计算机学院) Department of Computer Science and Engineering, Michigan State University(密歇根州立大学计算机科学与工程系) School of Information Science and Electronic Engineering, Shanghai Jiao Tong University(上海交通大学信息科学与电子工程学院) School of Computer Science, Shanghai Jiao Tong University(上海交通大学计算机科学学院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 MoGen通过区域语义锚和自适应多模态引导模块,实现了对多对象图像生成的可控性和灵活性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.14367 2026-01-12 cs.CV 57%

Hallucination Score: Towards Mitigating Hallucinations in Generative Image Super-Resolution

幻觉分数:朝着减轻生成图像超分辨率中的幻觉

Weiming Ren, Raghav Goyal, Zhiming Hu, Tristan Ty Aumentado-Armstrong, Iqbal Mohomed, Alex Levinshtein

机构 * University of Waterloo(多伦多大学) AI Center – Toronto, Samsung Electronics(三星电子人工智能中心)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 本文提出幻觉分数以衡量和减轻生成图像超分辨率中的幻觉问题,通过多模态大语言模型生成 HS 并用于模型微调。

Comments 31 pages, 21 figures, and 10 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.05397 2026-01-12 cs.CV 57%

Neural-Driven Image Editing

神经驱动的图像编辑

Pengfei Zhou, Jie Xia, Xiaopeng Peng, Wangbo Zhao, Zilong Ye, Zekai Li, Suorong Yang, Jiadong Pan, Yuanxiang Chen, Ziqiao Wang, Kai Wang, Qian Zheng, Hao Jin, Xiaojun Chang, Gang Pan, Shurong Dong, Kaipeng Zhang, Yang You

机构 * NUS(国立新加坡大学) ZJU(浙江大学) RIT(罗切斯特理工学院) NJU(南京大学) USTC(中国科学技术大学) Shanghai AI Lab(上海人工智能实验室) SII(上海信息所) Hangzhou RongNao Tech(杭州融纳科技)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 LoongX通过多模态神经信号驱动图像编辑,结合扩散模型和对比学习实现高性能图像编辑任务。

Comments 22 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏