arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2025-12-08 至 2025-12-08 共收录 39 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 图文多模态 6 篇

2512.05137 2025-12-08 cs.CV cs.AI 62%

ChromouVQA: Benchmarking Vision-Language Models under Chromatic Camouflaged Images

ChromouVQA:在色度伪装图像下评估视觉-语言模型的基准测试

Yunfei Zhang, Yizhuo He, Yuanxun Shao, Zhengtao Yao, Haoyan Xu, Junhao Dong, Zhen Yao, Zhikang Dong

机构 * Amazon(亚马逊公司) Google(谷歌公司) MurcuryMind(MurcuryMind公司) University of Southern California(南加州大学) Nanyang Technological University(南洋理工大学) Lehigh University(莱斯大学) Stony Brook University(石溪大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV、cs.AI

AI总结 ChromouVQA通过色度伪装图像评估视觉-语言模型在复杂背景下的表现,提出对比度配方提升形状恢复能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05482 2025-12-08 cs.CV 57%

Concept-based Explainable Data Mining with VLM for 3D Detection

基于概念的可解释数据挖掘与VLM用于3D检测

Mai Tsujimoto

机构 * The University of Tokyo(东京大学)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.CV

AI总结 本文提出基于概念的可解释数据挖掘方法,利用VLMs识别稀有物体以提升3D检测性能,减少标注负担并提高模型效果。

Comments 28 pages including appendix. Code: https://github.com/mm1129/concept_based_rare_detector_2025

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05145 2025-12-08 cs.CV 57%

Self-Improving VLM Judges Without Human Annotations

无需人工标注的自改进VLM评判模型

Inna Wanyin Lin, Yushi Hu, Shuyue Stella Li, Scott Geng, Pang Wei Koh, Luke Zettlemoyer, Tim Althoff, Marjan Ghazvininejad

机构 * FAIR at Meta(Meta 的 FAIR) University of Washington(华盛顿大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 无需人工标注,通过自训练提升VLM评判模型的准确性和多维度表现

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00713 2025-12-08 cs.CR cs.AI 57%

Concept-Guided Backdoor Attack on Vision Language Models

基于概念的视觉语言模型后门攻击

Haoyu Shen, Weimin Lyu, Haotian Xu, Tengfei Ma

机构 * Stony Brook University(石溪大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.AI

AI总结 本研究提出基于概念的视觉语言模型后门攻击方法,通过概念阈值污染和概念瓶颈模型引导未见后门两种技术,实现对模型生成文本的恶意替换。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.04555 2025-12-08 cs.RO cs.CV 57%

Evo-1: Lightweight Vision-Language-Action Model with Preserved Semantic Alignment

Evo-1:轻量级视觉-语言-动作模型,保持语义对齐

Tao Lin, Yilei Zhong, Yuxin Du, Jingjing Zhang, Jiting Liu, Yinxinyu Chen, Encheng Gu, Ziyan Liu, Hongyi Cai, Yanwen Zou, Lixing Zou, Zhaoye Zhou, Gen Li, Bo Zhao

机构 * School of AI, Shanghai Jiao Tong University(上海交通大学人工智能学院) EvoMind Tech(EvoMind科技) IAAR-Shanghai(IAAR-上海) SII Carnegie Mellon University(卡内基梅隆大学) University of Cambridge(剑桥大学) Nanyang Technological University(南洋理工大学)

专题命中 图文多模态 :multimodal(abstract);分类 cs.CV

AI总结 Evo-1是一种轻量级的视觉-语言-动作模型,通过减少计算并保持语义对齐,实现了高效的部署和强大的性能。

Comments Github: https://github.com/MINT-SJTU/Evo-1

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.10809 2025-12-08 cs.LG cs.AI 57%

Rethinking Sparse Autoencoders: Select-and-Project for Fairness and Control from Encoder Features Alone

重新思考稀疏自编码器:仅从编码器特征中进行选择和投影以实现公平性与控制

Antonio Bărbălau, Cristian Daniel Păduraru, Teodor Poncu, Alexandru Tifrea, Elena Burceanu

机构 * Bitdefender University Politehnica of Bucharest(巴特亚大学) ETH Zurich(苏黎世联邦理工学院)

专题命中 图文多模态 :cross-modal(abstract);分类 cs.AI

AI总结 本文提出了一种基于编码器特征的选择和投影框架,通过改进公平性和可控性,提升模型在视觉语言和大语言模型中的表现。

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 音频语音多模态 3 篇

2512.05745 2025-12-08 cs.CR cs.MM 83%

ARGUS: Defending Against Multimodal Indirect Prompt Injection via Steering Instruction-Following Behavior

ARGUS: 通过引导指令遵循行为防御多模态间接提示注入攻击

Weikai Lu, Ziqian Zeng, Kehua Zhang, Haoran Li, Huiping Zhuang, Ruidong Wang, Cen Chen, Hao Peng

专题命中 音频语音多模态 :multimodal(title,abstract);MLLM(abstract);分类 cs.MM

AI总结 ARGUS通过引导指令遵循行为,在表示空间中寻找最优防御方向,实现对多模态间接提示注入攻击的有效防御,同时保持模型性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05126 2025-12-08 eess.AS cs.AI cs.CL cs.CV cs.MM cs.SD 73%

SyncVoice: Towards Video Dubbing with Vision-Augmented Pretrained TTS Model

SyncVoice:面向视频配音的视觉增强预训练TTS模型

Kaidi Wang, Yi He, Wenhao Guan, Weijie Wu, Hongwu Ding, Xiong Zhang, Di Wu, Meng Meng, Jian Luan, Lin Li, Qingyang Hong

机构 * School of Informatics, Xiamen University, China(厦门大学信息学院) MiLM Plus, Xiaomi Inc., China(小米公司) School of Electronic Science and Engineering, Xiamen University, China(厦门大学电子科学与技术学院)

专题命中 音频语音多模态 :audio-visual(abstract);分类 cs.CV、cs.CL、cs.AI

AI总结 SyncVoice通过视觉增强预训练TTS模型,提升视频配音的语音自然度和音视频同步性能,适用于多语言视频翻译场景。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05508 2025-12-08 cs.SD cs.AI cs.LG 57%

Lyrics Matter: Exploiting the Power of Learnt Representations for Music Popularity Prediction

歌词很重要:利用学习到的表示力预测音乐流行度

Yash Choudhary, Preeti Rao, Pushpak Bhattacharyya

机构 * Indian Institute of Technology Bombay(印度理工学院班加罗尔) Department of Electrical Engineering, IIT Bombay(印度理工学院班加罗尔电气工程系) CFILT, Department of Computer Science, IIT Bombay(印度理工学院班加罗尔计算机科学系)

专题命中 音频语音多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文提出利用大语言模型提取歌词嵌入,结合音频和社交数据,提升音乐流行度预测精度,实验显示在SpotGenTrack数据集上MAE和MSE分别提升9%和20%。

Comments 8 pages

详情

展开后加载摘要…

URL PDF HTML 收藏

3. 视频多模态 4 篇

2412.07755 2025-12-08 cs.CV cs.AI cs.GR cs.RO 81%

SAT: Dynamic Spatial Aptitude Training for Multimodal Language Models

SAT:多模态语言模型的动态空间能力训练

Arijit Ray, Jiafei Duan, Ellis Brown, Reuben Tan, Dina Bashkirova, Rose Hendrix, Kiana Ehsani, Aniruddha Kembhavi, Bryan A. Plummer, Ranjay Krishna, Kuo-Hao Zeng, Kate Saenko

机构 * Boston University(波士顿大学) University of Washington(华盛顿大学) Allen Institute for AI(人工智能研究院) Microsoft Research(微软研究院) New York University(纽约大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 SAT通过模拟数据提升多模态语言模型在动态空间推理中的能力,实验表明其在多个基准测试中优于现有方法。

Comments Accepted to COLM 2025. Project webpage: https://arijitray.com/SAT/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.19907 2025-12-08 cs.CV 79%

MHB: Multimodal Handshape-aware Boundary Detection for Continuous Sign Language Recognition

MHB: 多模态手形感知边界检测用于连续手语识别

Mingyu Zhao, Zhanfu Yang, Yang Zhou, Zhaoyang Xia, Can Jin, Xiaoxiao He, Dimitris N. Metaxas

机构 * Rutgers University(罗杰斯大学)

专题命中 视频多模态 :multimodal(title,abstract);分类 cs.CV

AI总结 本文提出MHB方法,通过多模态融合结合3D骨骼特征和手形信息,提升连续手语识别的鲁棒性和准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.07978 2025-12-08 cs.CV 57%

Martian World Model: Controllable Video Synthesis with Physically Accurate 3D Reconstructions

火星世界模型:可控视频合成与物理准确的3D重建

Longfei Li, Zhiwen Fan, Wenyan Cong, Xinhang Liu, Yuyang Yin, Matt Foutter, Panwang Pan, Chenyu You, Yue Wang, Zhangyang Wang, Yao Zhao, Marco Pavone, Yunchao Wei

机构 * BJTU(北京工业大学) UT Austin(德克萨斯大学奥斯汀分校) HKUST(香港科技大学) Stanford University(斯坦福大学) XMU(厦门大学) SBU(雪城大学) USC(南加州大学) NVIDIA(英伟达)

专题命中 视频多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出M3arsSynth和MarsGen,通过物理准确的3D重建生成逼真的火星视频,提升任务模拟与机器人训练的可视化效果。

Comments Project Page: https://marsgenai.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.18473 2025-12-08 math.OC 50%

PDPO: Parametric Density Path Optimization

PDPO:参数密度路径优化

Sebastian Gutierrez Hernandez, Peng Chen, Haomin Zhou

专题命中 视频多模态 :multimodal(abstract)

AI总结 PDPO通过参数映射将无限维密度优化转化为有限维问题,有效解决多模态和高维路径优化问题,优于现有方法。

Comments 28 pages, 16 figures

Journal ref NeurIPS 2025

详情

展开后加载摘要…

URL PDF HTML 收藏

4. 跨模态检索 2 篇

2209.05463 2025-12-08 cs.CY cs.AI 79%

Modelling Business Agreements in the Multimodal Transportation Domain through Ontological Smart Contracts

通过本体智能合约建模多模态交通运输领域的商业协议

Mario Scrocca, Marco Comerio, Alessio Carenini, Irene Celino

专题命中 跨模态检索 :multimodal(title,abstract);分类 cs.AI

AI总结 本文提出通过本体智能合约建模多模式交通运输领域的商业协议,展示其在拼车场景中的应用及优势。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05908 2025-12-08 cs.SE cs.AI cs.CL cs.IR 62%

Natural Language Summarization Enables Multi-Repository Bug Localization by LLMs in Microservice Architectures

自然语言摘要通过LLM在微服务架构中多仓库Bug定位

Amirkia Rafiei Oskooei, S. Selcan Yukcu, Mehmet Cevheri Bozoglan, Mehmet S. Aktas

机构 * Yildiz Technical University, Dept. of Computer Eng.(Yildiz技术大学,计算机工程系) Yildiz Technical University, Department of Computer Engineering(Yildiz技术大学,计算机工程系)

专题命中 跨模态检索 :cross-modal(abstract);分类 cs.CL、cs.AI

AI总结 利用自然语言摘要和LLM技术,通过多仓库微服务架构实现更高效的Bug定位,显著提升定位准确率和透明度。

Comments Accepted at LLM4Code Workshop, ICSE 2026

详情

展开后加载摘要…

URL PDF HTML 收藏

5. 多模态生成 3 篇

2512.04563 2025-12-08 cs.CV 70%

COOPER: A Unified Model for Cooperative Perception and Reasoning in Spatial Intelligence

COOPER:一种用于空间智能中协作感知与推理的统一模型

Zefeng Zhang, Xiangzhao Hao, Hengzhu Tang, Zhenyu Zhang, Jiawei Sheng, Xiaodong Li, Zhenyang Li, Li Gao, Daiting Shi, Dawei Yin, Tingwen Liu

机构 * Institute of Information Engineering, Chinese Academy of Sciences(中国科学院信息工程研究所) School of Cyber Security, University of Chinese Academy of Sciences(中国科学院大学网络安全学院) Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所) Baidu Inc.(百度公司)

专题命中 多模态生成 :multimodal(abstract);MLLM(abstract);分类 cs.CV

AI总结 COOPER是一种统一的多模态大语言模型,通过整合深度和分割等辅助模态,提升空间感知与推理能力,实现空间智能的增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05965 2025-12-08 cs.CV 57%

EditThinker: Unlocking Iterative Reasoning for Any Image Editor

EditThinker: 解锁任何图像编辑器的迭代推理

Hongyu Li, Manyuan Zhang, Dian Zheng, Ziyu Guo, Yimeng Jia, Kaituo Feng, Hao Yu, Yexin Liu, Yan Feng, Peng Pei, Xunliang Cai, Linjiang Huang, Hongsheng Li, Si Liu

机构 * Beihang University(北航大学) Meituan(美团) CUHK MMLab CUHK IMIXR Tsinghua University(清华大学)

专题命中 多模态生成 :MLLM(abstract);分类 cs.CV

AI总结 EditThinker通过迭代推理框架提升图像编辑指令遵循能力,利用强化学习优化编辑过程,显著提高模型性能。

Comments Project page: https://appletea233.github.io/think-while-edit

详情

展开后加载摘要…

URL PDF HTML 收藏
2406.06211 2025-12-08 cs.CV 57%

iMotion-LLM: Instruction-Conditioned Trajectory Generation

iMotion-LLM:基于指令的轨迹生成

Abdulwahab Felemban, Nussair Hroub, Jian Ding, Eslam Abdelrahman, Xiaoqian Shen, Abduallah Mohamed, Mohamed Elhoseiny

机构 * King Abdullah University of Science and Technology (KAUST)(卡塔尔国王阿卜杜勒阿齐兹大学科学与技术学院) Meta Reality Labs(Meta现实实验室)

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 iMotion-LLM通过结合预训练LLM与轨迹预测模块,实现基于文本指令的交互式运动生成,提升自动驾驶中的轨迹生成准确性和安全性。

详情

展开后加载摘要…

URL PDF HTML 收藏

6. 多模态评测 10 篇

2511.16334 2025-12-08 cs.AI cs.CL 81%

OpenMMReasoner: Pushing the Frontiers for Multimodal Reasoning with an Open and General Recipe

OpenMMReasoner: 推动多模态推理的前沿研究:一个开放且通用的配方

Kaichen Zhang, Keming Wu, Zuhao Yang, Bo Li, Kairui Hu, Bin Wang, Ziwei Liu, Xingxuan Li, Lidong Bing

机构 * MiroMind AI Nanyang Technological University(南洋理工大学) Tsinghua University(清华大学) LMMs-Lab Team(多模态实验室团队)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL、cs.AI

AI总结 OpenMMReasoner提出了一种开放且通用的多模态推理训练配方,通过两阶段方法提升推理性能,实现在多个基准测试中超越现有基线。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05930 2025-12-08 cs.AI 79%

PRiSM: An Agentic Multimodal Benchmark for Scientific Reasoning via Python-Grounded Evaluation

PRiSM:一种基于Python基础评估的科学推理多模态基准

Shima Imani, Seungwhan Moon, Adel Ahmadyan, Lu Zhang, Kirmani Ahmed, Babak Damavandi

机构 * Meta Reality Lab(Meta现实实验室)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.AI

AI总结 PRiSM通过基于Python代码的动态多模态基准,评估科学推理能力,揭示VLMs在科学领域中的局限性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.02738 2025-12-08 cs.CV 79%

Open-PMC-18M: A High-Fidelity Large Scale Medical Dataset for Multimodal Representation Learning

Open-PMC-18M: 一种高保真大规模医学数据集用于多模态表示学习

Negin Baghbanzadeh, Mohammed Saidul Islam, Sajad Ashkezari, Elham Dolatabadi, Arash Afkanpour

机构 * Vector Institute(Vector研究所) York University(约克大学) University of Waterloo(滑铁卢大学)

专题命中 多模态评测 :multimodal(title);image-text(abstract);分类 cs.CV

AI总结 Open-PMC-18M通过高保真医学数据集提升多模态表示学习性能,实现子图提取与上下文丰富化,支持多种视觉-语言任务。

Comments 21 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2504.20454 2025-12-08 eess.IV cs.CV 79%

LymphAtlas- A Unified Multimodal Lymphoma Imaging Repository Delivering AI-Enhanced Diagnostic Insight

LymphAtlas- 一种统一的多模态淋巴瘤影像存储库,提供人工智能增强的诊断洞察

Jiajun Ding, Beiyao Zhu, Xiaosheng Liu, Lishen Zhang, Zhao Liu

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 LymphAtlas通过整合PET与CT数据,构建高质量多模态数据集,提升淋巴瘤分割精度与临床应用价值。

Comments 12 pages,3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05119 2025-12-08 cs.IR cs.AI cs.CL 73%

RAG-IGBench: Innovative Evaluation for RAG-based Interleaved Generation in Open-domain Question Answering

RAG-IGBench: 用于开放领域问答中基于检索增强生成的交错生成的创新评估

Rongyang Zhang, Yuqing Huang, Chengqiang Lu, Qimeng Wang, Yan Gao, Yi Wu, Yao Hu, Yin Xu, Wei Wang, Hao Wang, Enhong Chen

机构 * State Key Laboratory of Cognitive Intelligence, University of Science and Technology of China(认知智能国家重点实验室,中国科学技术大学) Xiaohongshu Inc.(小红书公司) Xi’an Jiaotong University(西安交通大学)

专题命中 多模态评测 :multimodal(abstract);image-text(abstract);分类 cs.CL、cs.AI

AI总结 RAG-IGBench通过创新的评估指标和多模态数据,评估基于检索增强生成的交错生成任务,验证了模型在开放领域问答中的性能提升。

Comments 26 pages, 6 figures, NeurIPS 2025 D&B Track poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05557 2025-12-08 cs.CV cs.AI 62%

2K-Characters-10K-Stories: A Quality-Gated Stylized Narrative Dataset with Disentangled Control and Sequence Consistency

2K角色-10K故事:一个具有解耦控制和序列一致性的质量门控风格化叙事数据集

Xingxi Yin, Yicheng Li, Gong Yan, Chenglin Li, Jian Zhao, Cong Huang, Yue Deng, Yin Zhang

机构 * Zhejiang University(浙江大学) Zhongguancun Institute of Artificial Intelligence(中关村人工智能研究院)

专题命中 多模态评测 :multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出2K-Characters-10K-Stories数据集,通过解耦控制和质量门控机制,实现高质量的风格化叙事生成。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05950 2025-12-08 cs.LG cs.AI 57%

Impugan: Learning Conditional Generative Models for Robust Data Imputation

Impugan:用于鲁棒数据填补的条件生成模型

Zalish Mahmud, Anantaa Kotal, Aritran Piplai

机构 * Computer Science(计算机科学) The University of Texas at El Paso(德克萨斯理工大学)

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 Impugan通过条件生成对抗网络实现鲁棒的数据填补与异质数据集整合,有效处理复杂数据中的缺失值问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.05122 2025-12-08 cs.AI 57%

Documenting SME Processes with Conversational AI: From Tacit Knowledge to BPMN

用对话式AI记录中小企业流程:从隐性知识到BPMN

Unnikrishnan Radhakrishnan

专题命中 多模态评测 :multimodal(abstract);分类 cs.AI

AI总结 本文提出一种基于LLM的对话助手,通过访谈式对话将中小企业隐性知识转化为BPMN图表,实现流程文档化,降低文档化门槛,提升运营透明度。

Comments Presented at 2025 International Workshop on Low-Cost Digital Solutions for Industrial Automation (LODISA)

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.15747 2025-12-08 cs.CV eess.IV 57%

A Strong View-Free Baseline Approach for Single-View Image Guided Point Cloud Completion

一种强视图无关的单视图图像引导点云补全基线方法

Fangzhou Lin, Zilin Dai, Rigved Sanku, Songlin Hou, Kazunori D Yamada, Haichong K. Zhang, Ziming Zhang

机构 * Department of Robotics Engineering, Worcester Polytechnic Institute(机器人工程系,沃斯特理工学院) Dell Technologies(戴尔技术) Graduate School of Information Sciences, Tohoku University(信息科学研究生院,东北大学) Department of Computer Science, Worcester Polytechnic Institute(计算机科学系,沃斯特理工学院) Harvard Kenneth C. Griffin Graduate School of Arts and Sciences(哈佛肯尼斯·C·格里芬艺术与科学研究生院) Department of Biomedical Engineering, Worcester Polytechnic Institute(生物医学工程系,沃斯特理工学院) Department of Electrical & Computer Engineering, Worcester Polytechnic Institute(电气与计算机工程系,沃斯特理工学院)

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出一种视图无关的单视图图像引导点云补全基线方法,通过多分支编码器-解码器网络和层次自融合机制,有效提升点云补全性能。

Comments 7 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.15436 2025-12-08 cs.CV 57%

Adaptive Chain-of-Focus Reasoning via Dynamic Visual Search and Zooming for Efficient VLMs

基于动态视觉搜索和缩放的自适应聚焦推理方法用于高效VLMs

Xintong Zhang, Zhi Gao, Bofei Zhang, Pengxiang Li, Xiaowen Zhang, Yang Liu, Tao Yuan, Yuwei Wu, Yunde Jia, Song-Chun Zhu, Qing Li

机构 * organization= School of Computer Science \& Technology, Beijing Institute of Technology , city= Beijing , country= China organization= State Key Laboratory of General Artificial Intelligence, BIGAI , city= Beijing , country= China organization= School of Intelligence Science Technology, Peking University , city= Beijing , country= China organization= Guangdong Laboratory of Machine Perception Intelligent Computing, Shenzhen MSU--BIT University , city= Shenzhen , country= China organization= Department of Automation, Tsinghua University , city= Beijing , country= China

专题命中 多模态评测 :multimodal(abstract);分类 cs.CV

AI总结 本文提出基于动态视觉搜索和缩放的自适应聚焦推理方法,提升VLMs的多模态推理效率和实际应用效果。

Comments https://github.com/xtong-zhang/Chain-of-Focus

详情

展开后加载摘要…

URL PDF HTML 收藏

7. 多模态Agent 2 篇

2512.05824 2025-12-08 cs.AI cs.CV 81%

Multimodal Oncology Agent for IDH1 Mutation Prediction in Low-Grade Glioma

多模态肿瘤代理用于低级别胶质瘤IDH1突变预测

Hafsa Akebli, Adam Shephard, Vincenzo Della Mea, Nasir Rajpoot

机构 * Department of Mathematics, Computer Science and Physics, University of Udine(乌迪大学数学、计算机科学与物理系) Tissue Image Analytics Centre, Department of Computer Science, University of Warwick(沃里克大学计算机科学系组织图像分析中心) Histofy Ltd(Histofy有限公司)

专题命中 多模态Agent :multimodal(title,abstract);分类 cs.CV、cs.AI

AI总结 本研究提出多模态肿瘤代理,结合组织学工具和外部生物医学资源,实现低级别胶质瘤IDH1突变的高精度预测。

Comments 4 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.22487 2025-12-08 eess.SY cs.SY 50%

A Multi-Objective Simultaneous Routing, Facility Location and Allocation Model for Earthquake Emergency Logistics

面向地震应急物流的多目标同时路由、设施选址与分配模型

Sakineh Khodadadi, Tohid Kargar Tasooji, Afshin Shariat-Mohayman, Navid Kalantari

专题命中 多模态Agent :multi-modal(abstract)

AI总结 本文提出多目标优化模型,同时优化地震应急物流中的路由、设施选址与医院分配,以减少需求未满足、伤员未服务和经济成本。

Comments One of the authors does not agree to publish the paper on arXiv

详情

展开后加载摘要…

URL PDF HTML 收藏