arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

共收录 4895 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 其他多模态 4895 篇

1607.02660 2016-07-12 cs.HC cs.AI cs.CV 62%

Augmenting Supervised Emotion Recognition with Rule-Based Decision Model

Amol Patwardhan, Gerald Knapp

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV、cs.AI

Comments 8 pages, 6 figures, 23 tables, IEEE TAC (in review)

详情

展开后加载摘要…

URL PDF HTML 收藏
1606.03333 2016-06-13 cs.MM cs.CL cs.IR 62%

Automatic Genre and Show Identification of Broadcast Media

Mortaza Doulaty, Oscar Saz, Raymond W. M. Ng, Thomas Hain

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CL、cs.MM

Comments Proc. of 17th Interspeech (2016), San Francisco, California, USA

详情

展开后加载摘要…

URL PDF HTML 收藏
1411.1784 2014-11-10 cs.LG cs.AI cs.CV stat.ML 62%

Conditional Generative Adversarial Nets

Mehdi Mirza, Simon Osindero

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1205.1365 2012-05-08 cs.MM cs.CV 62%

Image Enhancement with Statistical Estimation

Aroop Mukherjee, Soumen Kanrar

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV、cs.MM

Comments 9 pages,6 figures; ISSN:0975-5578 (Online); 0975-5934 (Print)

Journal ref The International Journal of Multimedia & Its Applications (IJMA) April 2012, Volume 4, Number 2, page 59-67

详情

展开后加载摘要…

URL PDF HTML 收藏
cs/0105003 2009-11-30 cs.CL cs.AI 62%

Rule Writing or Annotation: Cost-efficient Resource Usage for Base Noun Phrase Chunking

Grace Ngai, David Yarowsky

专题命中 其他多模态 :cross-modal(abstract);分类 cs.CL、cs.AI

Comments 9 pages, 4 figures, appeared in ACL2000

Journal ref Proceedings of the 38th Annual Meeting of the Association for Computational Linguistics, pages 117-125, Hong Kong (2000)

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.07878 2025-10-21 cs.LG cs.AI eess.SP q-bio.NC 61%

Comparative Analysis of Deep Learning Approaches for Harmful Brain Activity Detection Using EEG

Shivraj Singh Bhatti, Aryan Yadav, Mitali Monga, Neeraj Kumar

机构 * Department of Computer Science and Engineering(计算机科学与工程系) Thapar Institute of Engineering and Technology(泰帕尔工程与技术学院)

专题命中 其他多模态 :multimodal(abstract,comments);分类 cs.AI

Comments 6 pages, 5 figures. Presented at IEEE CICT 2024. The paper discusses the application of multimodal data and training strategies in EEG-based brain activity classification

Journal ref 2024 IEEE 8th Int. Conf. on Info. and Comm. Tech. (CICT), pp. 1-6

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.01542 2025-05-06 cs.HC cs.AI 61%

Emotions in the Loop: A Survey of Affective Computing for Emotional Support

Karishma Hegde, Hemadri Jayalath

机构 * School of Computing University of Georgia(计算学院 佐治亚大学)

专题命中 其他多模态 :multimodal(abstract,comments);分类 cs.AI

Comments 20 pages, 7 tables, 96 references. Survey paper on affective computing applications using large language models, multimodal AI, and therapeutic chatbots

详情

展开后加载摘要…

URL PDF HTML 收藏
2106.01536 2021-09-21 cs.CL 61%

BERT meets LIWC: Exploring State-of-the-Art Language Models for Predicting Communication Behavior in Couples' Conflict Interactions

Jacopo Biggiogera, George Boateng, Peter Hilpert, Matthew Vowels, Guy Bodenmann, Mona Neysari, Fridtjof Nussbeck, Tobias Kowatsch

专题命中 其他多模态 :multimodal(abstract,comments);分类 cs.CL

Comments 5 pages. Accepted at the 2nd Workshop on Social Affective Multimodal Interaction for Health (SAMIH) at ICMI 2021

详情

展开后加载摘要…

URL PDF HTML 收藏
1807.01466 2018-07-05 cs.CL 61%

Polarity and Intensity: the Two Aspects of Sentiment Analysis

Leimin Tian, Catherine Lai, Johanna D. Moore

专题命中 其他多模态 :multimodal(abstract,comments);分类 cs.CL

Comments Published at the First Grand Challenge and Workshop on Human Multimodal Language (Challenge-HML) of ACL 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.13577 2026-08-26 cs.CV cs.LG cs.RO 版本更新 57%

Adaptive Multi-Mode Out-of-Distribution Detection for Trajectory Prediction in Autonomous Vehicles

动态感知:面向自动驾驶轨迹预测的自适应多模式异常检测

Tongfei Guo, Lili Su

机构 * Department of Electrical and Computer Engineering, Northeastern University(东北大学电气与计算机工程系)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文提出一种自适应多模式异常检测框架,通过建模误差模式提升轨迹预测的鲁棒性,在复杂驾驶环境中实现更高效的异常检测。

Comments Accepted to the 2026 IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS 2026)

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22183 2026-08-25 cs.CV cs.IR cs.LG 新提交 57%

VERDICT: Agreement Beats Pixel-Space Verification in Real-Document OCSR

VERDICT:在真实文档的光学化学结构识别(OCSR)中,一致性优于像素空间验证

Yani Guan, Dengpan Dong, Shuang Luo, Zi Wei, Joah Han, Dan Hannah, Yumin Zhang, Qichao Hu, Kang Xu

机构 * SES AI Corporation(SES AI公司)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 该研究提出VERDICT方法,通过多架构识别器间的一致性而非像素空间验证,实现真实文档OCSR的可靠预测,在多数据集上表现优异,可构建经验证的分子数据库并用于分子记录检索。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.04930 2026-08-25 cs.LG cs.AI 版本更新 57%

SVI-DAG: A Structured Variational Inference Approach to Bayesian Causal Discovery

SVI-DAG:一种用于贝叶斯因果发现的结构化变分推断方法

Shrenik Zinage

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 SVI-DAG是一种贝叶斯因果发现的结构化变分推断方法,用归一化流建模边依赖、斯坦变分梯度下降优化,在不确定性量化上优于5种现有方法,结构准确性具竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.20423 2026-08-24 cs.LG cs.AI 新提交 57%

From Thermal Preference Prediction to Adaptive Thermal Intervention: A Reinforcement Learning Approach Using Physiological and Environmental Sensing

从热偏好预测到自适应热干预:一种使用生理与环境感知的强化学习方法

Isibor Kennedy Ihianle, Emmanuel Manu, Ehsan Asnaashari, Mojgan Jadidi, Pedro Machado, Amrit Sagoo, Ahmad Lotfi

机构 * Nottingham Trent University(诺丁汉特伦特大学) York University(约克大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 针对传统HVAC系统无法捕捉个体生理变异性的问题,提出整合多模态感知与强化学习的两阶段个性化热舒适方法,助力开发更响应的建筑控制策略。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.17447 2026-08-19 cs.CV 新提交 57%

NGS-Marker: Robust Native Watermarking for 3D Gaussian Splatting

NGS-Marker:面向3D高斯溅射(3DGS)的鲁棒原生水印

Hao Qin, Yukai Sun, Luyuan Chen, Mengxu Lu, Feng Zhang, Ming Kong, Zhenhong Du, Qiang Zhu

机构 * Zhejiang University(浙江大学) Zhejiang Key Laboratory of Geographic Information Science(浙江省地理信息科学重点实验室)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 该研究针对现有3D高斯溅射水印技术无法抵御部分侵权的问题,提出NGS-Marker原生水印框架,通过联合训练的注入器与解码器及梯度渐进注入策略实现全场景覆盖,可抵御部分侵权并支持混合与多模态水印,具备实际部署灵活性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.15115 2026-08-18 cs.CV 新提交 57%

Perspective-Invariant Attack with Enhanced Transferability of Adversarial Examples

具有增强对抗样本迁移性的视角不变攻击

Kaisheng Liang, Yiming Cao, Bin Xiao

机构 * The Hong Kong Polytechnic University(香港理工大学)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 针对对抗样本跨模型迁移性带来的安全威胁,提出视角不变攻击(PIA)及其扩展PIA-Mix,通过多自由度顶点采样策略提升对抗样本迁移性,实验显示其性能优于当前最优基于迁移的攻击方法。

Journal ref IEEE Transactions on Information Forensics and Security, vol. 21, pp. 6818-6831, 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.14306 2026-08-17 cs.AI cs.SY eess.SY 新提交 57%

Sensor-Driven Mission Synthesis for UAV/UGV Swarms: A TB-CSPN Coordination Architecture with Hardware-Enforced Safety

传感器驱动的无人机/无人车集群任务合成:具备硬件强制安全保障的TB-CSPN协同架构

Uwe M. Borghoff, Paolo Bottoni, Remo Pareschi

机构 * Institute for Software Technology, University of the Bundeswehr Munich(慕尼黑联邦国防军大学软件技术研究所) Department of Computer Science, Sapienza University of Rome(罗马大学计算机科学系) SofTware And Knowledge Engineering Lab(软件与知识工程实验室)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.AI

AI总结 本文提出一种具备硬件强制安全保障的TB-CSPN协同架构,用于异构无人机/无人车集群,结合多模态传感器观测,通过顾问与监督智能体实现可审计的决策路径,提升对抗环境下的韧性,经沿海监视案例验证其有效性。

Comments 15 pages, 4 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10434 2026-08-12 cs.AI 新提交 57%

Conversational versus Dashboard Explainable AI for UAV Intrusion Detection: An Empirical Study of Operator Trust and Reliance

无人机入侵检测中对话式与仪表盘式可解释人工智能:操作员信任与依赖的实证研究

Cong Chi Nguyen, Trang Mai Xuan, Vu-Duc Ngo, Kim-Ngan Thi Nguyen, Trong-Nghia Nguyen, Thien Van Luong

机构 * Phenikaa University(菲卡大学) Phenikaa School of Computing(菲卡计算机学院) MobiFone Corporation(MobiFone集团) MobiFone HighTech Center(MobiFone高科技中心) National Economics University(国民经济大学) College of Technology(技术学院)

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本研究对比对话式与仪表盘式XAI界面对无人机入侵检测操作员信任和依赖的影响,发现对话式界面可用性更高但易引发过度依赖,为未来XAI系统设计提供启示。

Comments 12 pages, 3 figures, EIDT conference

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.10203 2026-08-12 cs.CV 新提交 57%

A Convolutional Layer Activation Dimensionality Reduction for Out-of-Distribution and Adversarial Attack Detection Methods

用于分布外样本与对抗攻击检测方法的卷积层激活降维技术

Leandro de Souza Rosa, Lorenzo Capelli, Clara Nunes Barrancos, Mauro Mangia, Riccardo Rovatti

机构 * Alma Mater Studiorum Università di Bologna(博洛尼亚大学) Advanced Research Center on Electronic Systems “Ercole De Castro” (ARCES) - Alma Mater Studiorum Università di Bologna(博洛尼亚大学“埃尔科莱·德·卡斯特罗”电子系统高级研究中心)

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本文针对卷积层激活降维的不足,提出一种可控高压缩率的新型降维方法,扩展两种最先进的检测方法并在OOD与对抗攻击检测任务上验证,其性能更优且计算内存占用更低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.19125 2026-08-11 cs.LG cs.AI 版本更新 57%

Transformer Circuits Can Realize Clustering Algorithms

Transformer电路可实现聚类算法

Kenneth L. Clarkson, Lior Horesh, Takuya Ito, Charlotte Park, Parikshit Ram

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 研究表明Transformer电路可实现Lloyd算法等精确聚类算法,提出的k均值Transformer聚类算法泛化性强且质量优于Lloyd算法,改动后可衍生多种新型聚类算法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2203.10961 2026-08-11 cs.LG cs.AI 57%

Bike Sharing Demand Prediction based on Knowledge Sharing across Modes: A Graph-based Deep Learning Approach

Yuebing Liang, Guan Huang, Zhan Zhao

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06758 2026-08-10 cs.CL 新提交 57%

Stockmark-Nemotron-3-Nano-Omni-JapanDocReader: Structured Document Parsing via Capability Injection and Forgetting Control

Stockmark-Nemotron-3-Nano-Omni-JapanDocReader:基于能力注入与遗忘控制的结构化文档解析模型

Shi Chen, Hayato Aida, Makoto Morinaga, Shohei Tanaka, Kosuke Arima

机构 * Stockmark Inc.(斯托克马克公司)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CL

AI总结 本研究提出Stockmark-Nemotron-3-Nano-Omni-JapanDocReader模型,通过能力注入与遗忘控制优化结构化文档解析性能,结合混合SFT与DAPO式RL取得优于SFT的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05684 2026-08-07 cs.RO cs.AI cs.LG cs.SY eess.SY 新提交 57%

Nonvisual Classification of Ground-Condition by Artificial Proprioception in an Amoeba-Inspired Autonomous Walking Robot

变形虫启发的自主步行机器人中基于人工本体感觉的地面状况非视觉分类

Hyoto Yamaguchi, Zenji Yatabe, Seiya Kasai

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 该研究为变形虫启发的自主步行机器人,整合传感器与储备池计算实现人工本体感觉,完成地面状况非视觉分类,可依地面切换步态并分析传感器贡献。

Comments 5 pages, 7 figures, The paper has been submitted to IEEE SCIS ISIS 2026 for consideration

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.02955 2026-08-05 cs.HC cs.AI cs.CY eess.IV 新提交 57%

Chat Debugging: An Exploratory Study of Human-AI Collaboration to Debug Analog Circuits

聊天调试:人机协作调试模拟电路的探索性研究

John Hu, Andrew Ash

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 该研究通过分析本科生与开源大型语言模型(LLMs)协作调试模拟电路的聊天记录,揭示了人机协作调试的模式、LLMs的优势与不足及学生的技能短板,为优化人机协作调试提供了依据。

Comments This is the accepted version of a paper to be presented at the 2026 IEEE Frontiers in Education Conference (FIE). The final version will be available via IEEE Xplore

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01112 2026-08-04 cs.AI 新提交 57%

Fighting Fire with Fire: On the Feasibility of Protecting Exercises Against AI Cheating

以毒攻毒:利用对抗机器学习保护练习题抵御AI作弊的可行性研究

Tobias Braun, Jonas Grebe, Louis Rethfeld, Marcus Rohrbach

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本研究探究利用对抗机器学习,通过在多模态选择题视觉组件添加扰动引导AI作弊者给出固定错误答案,再以统计检验检测作弊,以保护教育练习题抵御AI作弊的可行性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.00743 2026-08-04 cs.CV 新提交 57%

LUT: Latent Utility Training for Visual Reasoning

LUT:用于视觉推理的隐式效用训练

Jiaxuan Kang, Siyu Chen, Mingda Li, Mingjie Liu, Tianyue Wang, Zhaoyang Wei, Yongheng Zhang, Yanchao Hao, Zheng Wei

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出仅用标准VQA对训练的LUT框架,通过轨迹和步骤层面的隐式效用优化,在感知密集型视觉推理基准上性能优于现有隐式推理方法,且标注成本更低。

详情

展开后加载摘要…

URL PDF HTML 收藏
2604.15622 2026-08-04 cs.CV cs.LG 版本更新 57%

AdaDINO: Context-Adaptive DINO-Distilled Vision Foundation Models for Efficient Open-Vocabulary Edge Inference

AdaVFM:通过LLM引导执行实现边缘智能的自适应视觉基础模型

Yiwei Zhao, Yi Zheng, Huapeng Su, Jieyu Lin, Stefano Ambrogio, Cijo Jose, Michael Ramamonjisoa, Patrick Labatut, Barbara De Salvo, Chiao Liu, Phillip B. Gibbons, Ziyun Li

机构 * Carnegie Mellon University(卡内基梅隆大学) Meta

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 本文提出AdaVFM,一种通过LLM引导执行实现边缘设备上语言对齐视觉基础模型高效推理的自适应框架,通过动态调整计算实现性能与效率的平衡。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.29278 2026-08-03 cs.CV 新提交 57%

Training-Free Entity-Level Few-Shot Segmentation of Remote Sensing Images with Advection Refinement

基于平流优化的遥感图像无训练实体级小样本分割

Xueting Bai, Huan Ni

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 该研究针对现有跨域小样本分割方法训练成本高、预测结果碎片化的问题,提出一种基于平流优化的无训练实体级遥感图像小样本分割框架,可提升 SAM3 的相关适应能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2605.00177 2026-08-03 cs.GR cs.CV 版本更新 57%

FieryGS: In-the-Wild Fire Synthesis with Physics-Integrated Gaussian Splatting

FieryGS: 在真实世界中实现火灾合成的物理集成高斯点云方法

Qianfan Shen, Ningxiao Tao, Qiyu Dai, Tianle Chen, Minghan Qin, Yongjie Zhang, Mengyu Chu, Wenzheng Chen, Baoquan Chen

机构 * School of EECS, Peking University(电子工程系,北京大学) School of Intelligence Science and Technology, Peking University(智能科学与技术学院,北京大学) Yuanpei College, Peking University(元培学院,北京大学) ByteDance Seed(字节跳动种子) Wangxuan Institute of Computer Technology, Peking University(王璇计算机技术研究所,北京大学) Beijing Academy of Artificial Intelligence, Beijing, China(北京人工智能研究院,北京,中国)

专题命中 其他多模态 :multimodal(abstract);分类 cs.CV

AI总结 FieryGS通过整合物理准确的燃烧模拟与渲染,实现真实世界3D场景中逼真的火灾合成,结合多模态大语言模型进行物理材料推理,提升火灾动态的可控性和真实性。

Comments ICLR 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.27428 2026-07-31 physics.med-ph cs.AI 新提交 57%

Rethinking Artificial Intelligence in Medical Imaging: Assumptions, Reality, and Reframing

重新思考医学影像中的人工智能:假设、现实与重构

Arman Rahmim, Nourhan Bayasi, Xiaoxiao Li, Babak Saboury, Fereshteh Yousefirizi

专题命中 其他多模态 :multimodal(abstract);分类 cs.AI

AI总结 本文指出医学影像AI研究转化不足源于结构性错位,明确六大错位维度并提出重构路径,最终愿景是开发与医生对齐、扩展而非替代临床判断的智能体AI。

Comments 10 pages, 2 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13174 2026-07-31 eess.IV cs.CV cs.LG 版本更新 57%

Scalable Drift Monitoring in Medical Imaging AI

医学影像AI中的可扩展漂移监测

Jameson Merkow, Felix J. Dorfner, Xiyu Yang, Alexander Ersoy, Giridhar Dasegowda, Mannudeep Kalra, Matthew P. Lungren, Christopher P. Bridge, Ivan Tarapov

专题命中 其他多模态 :multi-modal(abstract);分类 cs.CV

AI总结 本研究针对医学影像AI的模型漂移与可靠性问题,开发了基于CheXstray框架的增强型可扩展漂移监测框架MMC+,经真实世界数据验证可有效检测数据偏移并预警性能偏差,助力AI在临床场景的应用。

详情

展开后加载摘要…

URL PDF HTML 收藏