arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4975 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 4975 篇

2509.23589 2026-03-06 cs.AI cs.CV cs.LG 62%

BridgeDrive: Diffusion Bridge Policy for Closed-Loop Trajectory Planning in Autonomous Driving

BridgeDrive: 基于扩散桥的闭环轨迹规划扩散策略

Shu Liu, Wenlin Chen, Weihao Li, Zheng Wang, Lijin Yang, Jianing Huang, Yipin Zhang, Zhongzhan Huang, Ze Cheng, Hao Yang

机构 * Bosch (China) Investment Ltd(博世(中国)投资有限公司)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 BridgeDrive提出一种基于锚点的扩散桥策略,通过实时闭环轨迹规划提升自动驾驶的安全性和效率。

Comments Accepted for publication at ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.03657 2026-03-05 cs.CV cs.AI 62%

InEdit-Bench: Benchmarking Intermediate Logical Pathways for Intelligent Image Editing Models

InEdit-Bench:智能图像编辑模型中间逻辑路径的基准测试

Zhiqiang Sheng, Xumeng Han, Zhiwei Zhang, Zenghui Xiong, Yifan Ding, Aoxiang Ping, Xiang Li, Tong Guo, Yao Mao

机构 * State Key Laboratory of Optical Field Manipulation Science and Technology, Institute of Optics and Electronics, Chinese Academy of Sciences(光学场操控科学与技术国家重点实验室,光学电子研究所,中国科学院) National Laboratory on Adaptive Optics(自适应光学国家实验室) University of Chinese Academy of Sciences(中国科学院大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 InEdit-Bench通过评估图像编辑模型在中间逻辑路径推理中的表现,揭示了现有模型在动态推理和多步骤演变建模方面的不足,推动更智能的多模态生成模型发展。

Comments CVPR findings. Project page: https://github.com/SZStrong1/InEdit-Bench

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.14341 2026-03-03 cs.CV cs.AI cs.CY cs.LG 62%

Towards Transferable Defense Against Malicious Image Edits

面向恶意图像编辑的可迁移防御

Jie Zhang, Shuai Dong, Shiguang Shan, Xilin Chen

机构 * State Key Laboratory of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences (CAS)(人工智能安全国家重点实验室,计算技术研究所,中国科学院) University of China Academy of Sciences(中国科学院大学) School of Computer Science, China University of Geosciences(中国地质大学(武汉)计算机学院)

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

AI总结 TDAE通过双模优化提升图像对恶意编辑的免疫性,实现跨模型的可迁移防御。

Comments 14 pages, 5 figures, accepted by IEEE TPAMI

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.02253 2026-03-03 cs.CV cs.AI cs.LG 62%

DragFlow: Unleashing DiT Priors with Region Based Supervision for Drag Editing

DragFlow: 通过基于区域的监督释放DiT先验以实现拖拽编辑

Zihan Zhou, Shilin Lu, Shuli Leng, Shaocong Zhang, Zhuming Lian, Xinlei Yu, Adams Wai-Kin Kong

机构 * Nanyang Technological University(南洋理工大学) National University of Singapore(国立新加坡大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 DragFlow通过基于区域的监督利用FLUX先验,改进基于拖拽的图像编辑效果,实现对点式和区域式基线的超越。

Comments Accepted by ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16145 2026-03-03 cs.AI cs.CL 62%

SpiroLLM: Finetuning Pretrained LLMs to Understand Spirogram Time Series with Clinical Validation in COPD Reporting

SpiroLLM:通过临床验证在COPD报告中微调预训练大语言模型以理解肺功能时间序列

Shuhao Mei, Yongchao Long, Xiaoyu Xiao, Shan Cao, Xiaobo Han, Shijia Geng, Jinbo Sun, Yuxi Zhou, Shenda Hong

机构 * Guangzhou Institute of Technology, Xidian University, Xi’an, China(广州科技研究院,西安电子科技大学,中国) Department of Computer Science, Tianjin University of Technology, Tianjin, China(天津理工大学计算机学院,天津,中国) Department of Respiratory, The Second Hospital of Tianjin Medical University, China(天津医科大学第二医院呼吸科,中国) College of Pulmonary and Critical Care Medicine, Chinese PLA General Hospital, Beijing, China(中国人民解放军总医院呼吸与危重症医学科,北京,中国) HeartVoice Medical Technology, Hefei, China(合肥心声医疗技术有限公司,中国) School of Life Science and Technology, Xidian University, Xi’an, China(西安电子科技大学生命科学与技术学院,中国) National Institute of Health Data Science, Peking University, Beijing, China(北京大学国家健康数据科学研究院,北京,中国) Institute for Artificial Intelligence, Peking University, Beijing, China(北京大学人工智能研究院,北京,中国)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 SpiroLLM通过融合生理信号与大语言模型,实现对肺功能时间序列的解读,并在COPD诊断中展现出高准确性和稳健性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.12734 2026-03-03 cs.SD cs.AI cs.GR cs.HC eess.AS 62%

SounDiT: Geo-Contextual Soundscape-to-Landscape Generation

SounDiT:基于地理情境的声音景观到景观生成

Junbo Wang, Haofeng Tan, Bowen Liao, Albert Jiang, Teng Fei, Qixing Huang, Bing Zhou, Zhengzhong Tu, Shan Ye, Yuhao Kang

机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) University of Tennessee, Knoxville(田纳西大学基洛纳分校) University of South Carolina(南卡罗来纳大学) Arizona State University(亚利桑那州立大学) University of Canterbury(坎特伯雷大学) Texas A&M University(德克萨斯A&M大学) University of Wisconsin-Madison(威斯康星大学麦迪逊分校)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.AI、eess.AS

AI总结 SounDiT通过结合环境声音景观和地理情境条件,生成地理上一致的景观图像,并引入Place Similarity Score评估生成一致性。

Comments 12 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2603.00155 2026-03-03 cs.CV cs.AI cs.IR 62%

EfficientPosterGen: Semantic-aware Efficient Poster Generation via Token Compression and Accurate Violation Detection

EfficientPosterGen: 通过令牌压缩和准确违规检测的语义感知高效海报生成

Wenxin Tang, Jingyu Xiao, Yanpei Gong, Fengyuan Ran, Tongchuan Xia, Junliang Liu, Man Ho Lam, Wenxuan Wang, Michael R. Lyu

机构 * Tsinghua University(清华大学) The Chinese University of Hong Kong(香港中文大学) Harbin Institute of Technology(哈尔滨工业大学) Wuhan University(武汉大学) Beijing University of Posts and Telecommunications(北京邮电大学) Dalian Maritime University(大连海事大学) Renmin University of China(中国人民大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 EfficientPosterGen通过语义感知检索、视觉上下文压缩和无代理布局检测技术,实现高效且可靠的学术海报自动生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.23969 2026-03-02 cs.MM cs.CV 62%

MSVBench: Towards Human-Level Evaluation of Multi-Shot Video Generation

MSVBench: 向多镜头视频生成的人机水平评估迈进

Haoyuan Shi, Yunxin Li, Nanhao Deng, Zhenran Xu, Xinyu Chen, Longyue Wang, Baotian Hu, Min Zhang

机构 * Harbin Institute of Technology (Shenzhen)(哈尔滨工业大学(深圳)) Alibaba International Digital Commerce(阿里巴巴国际数字商业)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.MM

AI总结 MSVBench通过引入分层脚本和参考图像,提出混合评估框架,验证了视频生成模型的连贯性和吸引力,并展示了其在多镜头视频生成中的有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.22624 2026-02-27 cs.CV cs.AI 62%

Instruction-based Image Editing with Planning, Reasoning, and Generation

基于指令的图像编辑与规划、推理和生成

Liya Ji, Chenyang Qi, Qifeng Chen

机构 * HKUST(香港科技大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出一种多模态模型,通过链式思考规划、编辑区域推理和编辑,提升基于指令的图像编辑能力,以应对更复杂的真实场景。

Comments 10 pages, 7 figures

Journal ref Proceedings of the IEEE/CVF International Conference on Computer Vision, 2025, Page 17506--17515

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.18022 2026-02-26 cs.CV cs.AI 62%

Dual-Channel Attention Guidance for Training-Free Image Editing Control in Diffusion Transformers

双通道注意力引导用于扩散变换器中无需训练的图像编辑控制

Guandong Li

机构 * iFLYTEK

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出双通道注意力引导方法,通过同时操控键通道和值通道实现无需训练的图像编辑控制,显著提升编辑保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.20520 2026-02-25 cs.CV cs.AI 62%

How Do Inpainting Artifacts Propagate to Language?

图像修复伪影如何传播到语言?

Pratham Yashwante, Davit Abrahamyan, Shresth Grover, Sukruth Rao

机构 * UC San Diego(圣迭戈大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 研究图像修复伪影对视觉-语言模型语言生成的影响,通过两阶段诊断方法分析重建保真度与描述质量的关系。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02437 2026-02-23 cs.CV cs.AI 62%

UniReason 1.0: A Unified Reasoning Framework for World Knowledge Aligned Image Generation and Editing

UniReason 1.0: 一个统一的推理框架,用于世界知识对齐的图像生成与编辑

Dianyi Wang, Chaofan Ma, Feng Han, Size Wu, Wei Song, Yibin Wang, Zhixiong Zhang, Tianhang Wang, Siyuan Wang, Zhongyu Wei, Jiaqi Wang

机构 * Fudan University(复旦大学) Shanghai Innovation Institute(上海创新研究院) Shanghai Jiao Tong University(上海交通大学) Nanyang Technological University(南洋理工大学) Zhejiang University(浙江大学) University of Southern California(美国南加州大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 UniReason 1.0通过统一推理框架整合图像生成与编辑,利用世界知识增强文本推理和视觉优化,提升复杂合成任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.11096 2026-02-12 cs.CL cs.AI 62%

Safety Recovery in Reasoning Models Is Only a Few Early Steering Steps Away

推理模型中的安全恢复仅需几步早期引导步骤

Soumya Suvra Ghosal, Souradip Chakraborty, Vaibhav Singh, Furong Huang, Dinesh Manocha, Amrit Singh Bedi

机构 * University of Maryland, College Park(马里兰大学哥伦比亚学院) IIT, Bombay(孟买印度理工学院) University of Central Florida(佛罗里达中央大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 SafeThink通过在推理早期干预减少攻击成功率,提升推理模型的安全性

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21963 2026-02-12 cs.CY cs.AI cs.CL cs.SI 62%

Industrialized Deception: The Collateral Effects of LLM-Generated Misinformation on Digital Ecosystems

工业化的欺骗:大语言模型生成的虚假信息对数字生态系统的影响

Alexander Loth, Martin Kappes, Marc-Oliver Pahl

机构 * Frankfurt University of Applied Sciences(法兰克福应用科学大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出JudgeGPT和RogueGPT工具,研究人类对AI生成虚假信息的感知与检测,探讨生成与检测之间的竞争及缓解策略。

Comments Accepted at ACM TheWebConf '26 Companion

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07983 2026-02-10 cs.AI cs.CL 62%

Accelerating Social Science Research via Agentic Hypothesization and Experimentation

通过代理假设和实验加速社会科学研究

Jishu Sen Gupta, Harini SI, Somesh Kumar Singh, Syed Mohamad Tawseeq, Yaman Kumar Singla, David Doermann, Rajiv Ratn Shah, Balaji Krishnamurthy

机构 * Adobe Media and Data Science Research (MDSR)(Adobe媒体与数据科学研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 EXPERIGEN通过代理框架实现端到端的科学发现,发现更多显著且预测性强的假设,并通过A/B测试验证其有效性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.08785 2026-02-10 cs.CV cs.AI 62%

DeltaSpace: A Semantic-aligned Feature Space for Flexible Text-guided Image Editing

DeltaSpace: 一种语义对齐的特征空间用于灵活的文本引导图像编辑

Yueming Lyu, Kang Zhao, Bo Peng, Huafeng Chen, Yue Jiang, Yingya Zhang, Jing Dong, Caifeng Shan

专题命中 多模态生成 :image-text(abstract);分类 cs.CV、cs.AI

AI总结 DeltaSpace通过语义对齐的特征空间实现文本引导图像编辑的灵活训练和推理,支持零样本推理和无需文本的训练。

Comments 18 pages. arXiv admin note: text overlap with arXiv:2303.06285

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.07176 2026-02-10 cs.CL cs.AI cs.ET cs.HC 62%

Open TutorAI: An Open-source Platform for Personalized and Immersive Learning with Generative AI

Open TutorAI: 一个基于生成AI的开源平台,用于个性化和沉浸式学习

Mohamed El Hajji, Tarek Ait Baha, Aicha Dakir, Hammou Fadili, Youssef Es-Saady

机构 * IRF-SIC Laboratory, Ibnou Zohr University(IRF-SIC实验室,伊本·扎赫尔大学) Regional Center for Education(教育与培训专业地区中心) Polydisciplinary Faculty of Taroudant, Ibnou Zohr University(塔鲁旦多学科学院,伊本·扎赫尔大学) Higher School of Technology of Guelmim, Ibnou Zohr University(盖尔米姆技术高等学校,伊本·扎赫尔大学) Paragraphe laboratory, Paris 8 and CY Cergy Paris Universities(Paragraphe实验室,巴黎8大学和CY塞克巴黎大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 Open TutorAI 是一个基于生成AI的开源平台,通过个性化和沉浸式学习体验提升教育效果。

Comments 19 pages, 15 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.02276 2026-02-06 cs.LG cs.AI cs.CL q-bio.QM 62%

CellForge: Agentic Design of Virtual Cell Models

CellForge: 虚拟细胞模型的代理设计

Xiangru Tang, Zhuoyun Yu, Jiapeng Chen, Yan Cui, Daniel Shao, Weixu Wang, Fang Wu, Yuchen Zhuang, Wenqi Shi, Zhi Huang, Arman Cohan, Xihong Lin, Fabian Theis, Smita Krishnaswamy, Mark Gerstein

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 CellForge通过多代理协作自主设计虚拟细胞模型,生成高竞争力的计算方法,推动计算生物学的自主科学方法开发。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.02820 2026-02-04 cs.LG cs.AI cs.CV 62%

From Tokens to Numbers: Continuous Number Modeling for SVG Generation

从标记到数字:用于SVG生成的连续数字建模

Michael Ogezi, Martin Bell, Freda Shi, Ethan Smith

机构 * Cheriton School of Computer Science, University of Waterloo, Waterloo, ON, Canada(滑铁卢大学计算机科学学院) Vector Institute(向量研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出连续数字建模(CNM)方法,通过直接建模连续数值提升SVG生成的效率与质量,实现训练速度提升30%及更高的视觉保真度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2602.01193 2026-02-03 cs.CL cs.CV 62%

Bridging Lexical Ambiguity and Vision: A Mini Review on Visual Word Sense Disambiguation

弥合词汇歧义与视觉:关于视觉词义消歧的简要综述

Shashini Nilukshi, Deshan Sumanathilaka

机构 * School of Computing Informatics(计算与信息学学院) Institute of Technology(技术研究所)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.CL

AI总结 本文综述了视觉词义消歧的发展,探讨了对比模型和LLM在解决词汇歧义中的作用,并指出未来发展方向。

Comments 2 figures, 2 Tables, Accepted at IEEE TIC 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20911 2026-01-30 cs.CV cs.AI 62%

Non-Markov Multi-Round Conversational Image Generation with History-Conditioned MLLMs

非马尔可夫多轮对话图像生成与历史条件化大语言模型

Haochen Zhang, Animesh Sinha, Felix Juefei-Xu, Haoyu Ma, Kunpeng Li, Zhipeng Fan, Meng Dong, Xiaoliang Dai, Tingbo Hou, Peizhao Zhang, Zecheng He

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本研究提出非马尔可夫多轮对话图像生成方法,通过历史条件化框架和数据构建策略提升多轮一致性与指令遵循性,同时保持单轮编辑能力。

Comments 19 pages, 19 figures, plan for TIP

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.03147 2026-01-30 cs.HC cs.CL cs.LG cs.SD eess.AS 62%

A conversational gesture synthesis system based on emotions and semantics

基于情感和语义的对话手势合成系统

Thanh Hoang-Minh

机构 * Department of Information Technology, VNUHCM -- University of Science(越南科学大学信息科技系)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、eess.AS

AI总结 本文提出DeepGesture,一种基于扩散的 gesture 合成系统,通过多模态信号生成具有情感和语义条件的手势,提升数字人类的自然表达能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.19613 2026-01-28 cs.CL cs.AI 62%

Up to 36x Speedup: Mask-based Parallel Inference Paradigm for Key Information Extraction in MLLMs

最高36倍提速:面向MLLMs关键信息提取的基于掩码的并行推断范式

Xinzhong Wang, Ya Guo, Jing Li, Huan Chen, Yi Tu, Yijie Hong, Gongshen Liu, Huijia Zhu

机构 * Shanghai Jiao Tong University(上海交通大学) Ant Info Security Lab, Ant Group(蚂蚁集团信息安全部实验室) Inner Mongolia Research Institute, Shanghai Jiao Tong University(内蒙古研究院,上海交通大学)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CL、cs.AI

AI总结 本文提出基于掩码的并行推断范式,通过并行生成提升MLLMs关键信息提取的效率,实现最高36倍的提速。

Comments Accepted by ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.12529 2026-01-28 cs.HC cs.AI cs.CL 62%

Accepted with Minor Revisions: Value of AI-Assisted Scientific Writing

接受修改后:人工智能辅助科学写作的价值

Sanchaita Hazra, Doeun Lee, Bodhisattwa Prasad Majumder, Sachin Kumar

机构 * The University of Utah(犹他大学) The Ohio State University(俄亥俄州立大学) Allen Institute for AI(人工智能研究院)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 研究探讨了AI辅助科学写作的有效性,发现AI生成摘要在披露来源信息后可达到与人工摘要相当的可接受性,且作者编辑行为受对AI作者身份的感知驱动。

Comments Published in ACM IUI 2026 (Paphos, Cyprus)

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.18234 2026-01-27 cs.CY cs.AI cs.CL 62%

Generative AI in Saudi Arabia: A National Survey of Adoption, Risks, and Public Perceptions

生成式人工智能在沙特阿拉伯:国家调查:采用、风险和公众认知

Abdulaziz AlDakheel, Ali Alshehre, Esraa Alamoudi, Moslim AlKhabbaz, Ahmed Aljohani, Raed Alharbi

机构 * College of Computing and Informatics, Saudi Electronic University(计算机与信息学院,沙特电子大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本研究通过全国调查分析沙特阿拉伯GenAI的采用情况、风险和公众认知,发现93%的受访者积极使用GenAI进行文本任务,但整体意识和理解不均衡,需加强AI素养和伦理培训。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17673 2026-01-27 cs.CV cs.AI 62%

Uni-RS: A Spatially Faithful Unified Understanding and Generation Model for Remote Sensing

Uni-RS: 一种用于遥感的具有空间忠实性的统一理解和生成模型

Weiyu Zhang, Yuan Hu, Yong Li, Yu Liu

机构 * Institute of Remote Sensing and Geographic Information System, School of Earth and Space Sciences, Peking University(遥感与地理信息系统研究所,地球与空间科学学院,北京大学) Department of Civil and Environmental Engineering, The Hong Kong University of Science and Technology(土木与环境工程系,香港科学与技术大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Uni-RS通过空间布局规划、空间感知查询监督和图像描述空间布局变化,提升遥感文本到图像生成的空间忠实性,同时保持多模态理解任务的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17096 2026-01-27 cs.CY cs.AI cs.CL 62%

Beyond Instrumental and Substitutive Paradigms: Introducing Machine Culture as an Emergent Phenomenon in Large Language Models

超越工具性和替代性范式:引入机器文化作为大型语言模型中的涌现现象

Yueqing Hu, Xinyang Peng, Yukun Zhao, Lin Qiu, Ka-lai Hung, Kaiping Peng

专题命中 多模态生成 :multimodal(abstract);分类 cs.CL、cs.AI

AI总结 本研究提出机器文化作为大型语言模型中的一种新兴现象,挑战传统工具性和替代性范式,揭示模型在文化表现上的独特特性。

Comments 16 pages, 6 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.17027 2026-01-27 cs.CV cs.AI 62%

Scientific Image Synthesis: Benchmarking, Methodologies, and Downstream Utility

科学图像合成:基准测试、方法论与下游应用

Honglin Lin, Chonghan Qin, Zheng Liu, Qizhi Pei, Yu Li, Zhanping Zhong, Xin Gao, Yanfeng Wang, Conghui He, Lijun Wu

机构 * Shanghai Jiao Tong University(上海交通大学) OpenDataLab, Shanghai Artificial Intelligence Laboratory(OpenDataLab,上海人工智能实验室) The University of Hong Kong(香港大学) Peking University(北京大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出ImgCoder框架和SciGenBench基准,通过逻辑驱动方法提升科学图像生成的结构精度,并展示微调LMMs在科学图像上的效果,验证了高保真合成在多模态推理中的潜力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.16007 2026-01-23 cs.CV cs.AI 62%

PhysicsMind: Sim and Real Mechanics Benchmarking for Physical Reasoning and Prediction in Foundational VLMs and World Models

PhysicsMind: 为基础多模态大语言模型和世界模型中的物理推理和预测进行仿真与现实力学基准测试

Chak-Wing Mak, Guanyu Zhu, Boyi Zhang, Hongji Li, Xiaowei Chi, Kevin Zhang, Yichen Wu, Yangfan He, Chun-Kai Fan, Wentao Lu, Kuangzhi Ge, Xinyu Fang, Hongyang He, Kuan Lu, Tianxiang Xu, Li Zhang, Yongxin Ni, Youhua Li, Shanghang Zhang

机构 * Peking University(北京大学) Mohamed bin Zayed University of Artificial Intelligence(莫扎伊德大学人工智能学院) National University of Singapore(新加坡国立大学) University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) University of Science and Technology of China(中国科学技术大学) Cornell University(康奈尔大学) Hong Kong Polytechnic University(香港理工大学) City University of Hong Kong(香港城市大学)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 PhysicsMind是一个结合现实和仿真环境的统一基准,用于评估基础多模态大语言模型和世界模型在物理推理和预测中的能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15664 2026-01-23 cs.CV cs.AI 62%

Skywork UniPic 3.0: Unified Multi-Image Composition via Sequence Modeling

Skywork UniPic 3.0:通过序列建模实现统一的多图像合成

Hongyang Wei, Hongbo Liu, Zidong Wang, Yi Peng, Baixin Xu, Size Wu, Xuying Zhang, Xianglong He, Zexiang Liu, Peiyu Wang, Xuchen Song, Yangguang Li, Yang Liu, Yahui Zhou

机构 * Skywork

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 Skywork UniPic 3.0通过序列建模实现统一的多图像合成,采用新颖的训练范式和高效的数据流程,在单图像编辑和多图像合成任务中均取得优异性能。

详情

展开后加载摘要…

URL PDF HTML 收藏