arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

多模态大模型

跨文本、图像、视频、音频等模态的大模型与学习方法。

2026-01-21 至 2026-01-21 共收录 151 信号源:cs.CV, cs.CL, cs.AI, cs.MM, eess.AS

1. 多模态生成 19 篇

2601.12326 2026-01-21 cs.CV 57%

EmoKGEdit: Training-free Affective Injection via Visual Cue Transformation

EmoKGEdit: 无训练情感注入 via 视觉线索转换

Jing Zhang, Bingjie Fan

专题命中 多模态生成 :multimodal(abstract);分类 cs.CV

AI总结 EmoKGEdit通过构建多模态情感关联知识图谱,实现无训练的精确情感注入,有效分离情感属性与布局特征,提升图像情感编辑的保真度和结构一致性。

Comments 11pages,10figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.11074 2026-01-21 cs.CV 57%

Evaluating Latent Generative Paradigms for High-Fidelity 3D Shape Completion from a Single Depth Image

评估用于从单个深度图像生成高保真3D形状补全的潜在生成范式

Matthias Humt, Ulrich Hillenbrand, Rudolph Triebel

机构 * German Aerospace Center(德国航空航天中心) TU Munich(慕尼黑技术大学) Karlsruhe Institute of Technology(卡尔斯鲁厄理工学院)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 本文比较了扩散模型和自回归模型在单个深度图像生成高保真3D形状补全任务中的性能,发现扩散模型在连续潜在空间中表现更优,而自回归模型在离散潜在空间中可匹敌或超越扩散模型。

Comments 16 pages, 4 figures, 19 tables. To appear in 3DV 2026. Project page: https://hummat.github.io/2026-3dv-genz/

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.07260 2026-01-21 cs.AI cs.LG 57%

PADiff: Predictive and Adaptive Diffusion Policies for Ad Hoc Teamwork

PADiff: 预测性与自适应扩散策略用于即兴团队合作

Hohei Chan, Xinzhi Zhang, Antao Xiang, Weinan Zhang, Mengchen Zhao

专题命中 多模态生成 :multimodal(abstract);分类 cs.AI

AI总结 PADiff通过整合队友预测信息,提升在非平稳即兴团队合作场景中的预测与适应能力,实现多模态协作模式的多样化。

Comments Accepted by AAAI 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.16869 2026-01-21 cs.GR cs.CV 57%

Controllable Video Generation: A Survey

可控视频生成:综述

Yue Ma, Kunyu Feng, Zhongyuan Hu, Xinyu Wang, Yucheng Wang, Mingzhe Zheng, Bingyuan Wang, Qinghe Wang, Xuanhua He, Hongfa Wang, Chenyang Zhu, Hongyu Liu, Yingqing He, Zeyu Wang, Zhifeng Li, Xiu Li, Sirui Han, Yike Guo, Wei Liu, Dan Xu, Linfeng Zhang, Qifeng Chen

机构 * Hong Kong University of Science and Technology(香港科技大学) Hong Kong University of Science and Technology(Guang Zhou)(香港科技大学(广州)) Tsinghua University(清华大学) Dalian University of Technology(大连理工大学) Tencent(腾讯)

专题命中 多模态生成 :multi-modal(abstract);分类 cs.CV

AI总结 本文综述了可控视频生成领域,探讨了视频扩散模型中的控制机制及不同控制信号类型的方法分类。

Comments project page: https://github.com/mayuelala/Awesome-Controllable-Video-Generation

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14228 2026-01-21 cs.LG 50%

Attention-Based Offline Reinforcement Learning and Clustering for Interpretable Sepsis Treatment

基于注意力机制的离线强化学习与聚类用于可解释的脓毒症治疗

Punit Kumar, Vaibhav Saran, Divyesh Patel, Nitin Kulkarni, Alina Vereshchaka

机构 * Department of Computer Science(计算机科学系) Engineering University at Buffalo Buffalo, New York, USA(布法罗大学工程学院)

专题命中 多模态生成 :multi-modal(abstract)

AI总结 本文提出基于注意力机制的离线强化学习与聚类方法,用于可解释的脓毒症治疗决策支持,通过多模块整合提升治疗准确性和可解释性。

Comments 8 pages, 6 figures, Conference: IEEE International Conference on Machine Learning and Applications 2025 (ICMLA 2025): https://www.icmla-conference.org/icmla25/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13250 2026-01-21 cs.RO 50%

Diffusion-based Inverse Model of a Distributed Tactile Sensor for Object Pose Estimation

基于扩散的分布式触觉传感器逆模型用于物体姿态估计

Ante Marić, Giammarco Caroleo, Alessandro Albini, Julius Jankowski, Perla Maiolino, Sylvain Calinon

机构 * Idiap Research Institute(Idiap研究机构) École Polytechnique Fédérale de Lausanne(瑞士联邦理工学院) Oxford Robotics Institute(牛津机器人研究所)

专题命中 多模态生成 :multimodal(abstract)

AI总结 本文提出基于扩散模型的分布式触觉传感器逆模型,用于在无视觉数据情况下提高物体姿态估计的精度和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12901 2026-01-21 cs.RO 50%

PlannerRFT: Reinforcing Diffusion Planners through Closed-Loop and Sample-Efficient Fine-Tuning

PlannerRFT: 通过闭环和样本高效微调强化扩散规划器

Hongchen Li, Tianyu Li, Jiazhi Yang, Haochen Tian, Caojun Wang, Lei Shi, Mingyang Shang, Zengrong Lin, Gaoqiang Wu, Zhihui Hao, Xianpeng Lang, Jia Hu, Hongyang Li

机构 * Tongji University(同济大学) Shanghai Innovation Institute(上海创新研究院) OpenDriveLab at The University of Hong Kong(香港大学OpenDrive实验室) Meituan(美团) Li Auto Inc.(李自动公司)

专题命中 多模态生成 :multi-modal(abstract)

AI总结 PlannerRFT通过闭环和样本高效微调提升扩散规划器的多模态轨迹生成能力,实现高效探索与鲁棒性增强。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20339 2026-01-21 cs.SD 50%

MMEDIT: A Unified Framework for Multi-Type Audio Editing via Audio Language Model

MMEDIT: 一种基于音频语言模型的多类型音频编辑统一框架

Ye Tao, Wen Wu, Chao Zhang, Mengyue Wu, Shuai Wang, Xuenan Xu

机构 * MoE Key Lab of Artificial Intelligence, X-LANCE Lab, Shanghai Jiao Tong University(人工智能莫埃实验室、X-LANCE实验室、上海交通大学) Shanghai AI Laboratory(上海人工智能实验室) Nanjing University(南京大学)

专题命中 多模态生成 :cross-modal(abstract)

AI总结 MMEdit提出一种基于音频语言模型的统一框架,通过扩展任务定义和设计数据合成管道,实现多类型音频编辑的高精度与高保真度。

Comments Under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24555 2026-01-21 cs.LG 50%

From Perception to Punchline: Empowering VLM with the Art of In-the-wild Meme

从感知到 punchline:通过野生表情包艺术赋能 VLM

Xueyan Li, Yingyi Xue, Mengjie Jiang, Qingzi Zhu, Yazhe Niu

机构 * Shanghai Artificial Intelligence Laboratory(上海人工智能实验室) School of Software Engineering(软件工程学院) Columbia Engineering(哥伦比亚工程学院) Columbia University(哥伦比亚大学) The Chinese University of Hong Kong MMLab(香港中文大学MMLab)

专题命中 多模态生成 :multimodal(abstract)

AI总结 HUMOR 通过分层推理和群体偏好对齐,提升 VLM 在多模态生成中的推理多样性与幽默质量。

Comments 46 pages, 20 figures

详情

展开后加载摘要…

URL PDF HTML 收藏

2. 多模态评测 45 篇

2601.12346 2026-01-21 cs.CV 83%

MMDeepResearch-Bench: A Benchmark for Multimodal Deep Research Agents

MMDeepResearch-Bench: 一个多模态深度研究代理的基准

Peizhou Huang, Zixuan Zhong, Zhongwei Wan, Donghao Zhou, Samiul Alam, Xin Wang, Zexin Li, Zhihao Dou, Li Zhu, Jing Xiong, Chaofan Tao, Yan Xu, Dimitrios Dimitriadis, Tuo Zhang, Mi Zhang

机构 * OSU(俄亥俄州立大学) Amazon(亚马逊公司) UMich(密歇根大学) UCL(伦敦大学学院) CUHK(香港中文大学) UCR(加州大学尔湾分校) CWRU(克里夫兰医学中心) HKU(香港大学)

专题命中 多模态评测 :multimodal(title,abstract);image-text(abstract);分类 cs.CV

AI总结 MMDeepResearch-Bench提出一个多模态深度研究代理的基准,强调报告式合成与引用证据的结合,揭示多模态完整性对深度研究代理的重要性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2410.13532 2026-01-21 cs.CV 83%

RemoteDet-Mamba: A Hybrid Mamba-CNN Network for Multi-modal Object Detection in Remote Sensing Images

RemoteDet-Mamba: 一种用于遥感图像多模态目标检测的混合Mamba-CNN网络

Kejun Ren, Xin Wu, Lianming Xu, Li Wang

机构 * School of Computer Science, Beijing University of Posts and Telecommunications, Beijing, China(计算机科学学院,北京邮电大学) School of Electronic Engineering, Beijing University of Posts and Telecommunications, Beijing, China(电子工程学院,北京邮电大学)

专题命中 多模态评测 :multi-modal(title,abstract);cross-modal(abstract);分类 cs.CV

AI总结 RemoteDet-Mamba通过融合多模态信息提升遥感图像中小目标的检测性能,采用轻量级融合机制降低计算复杂度,实验证明其在检测精度和效率上的优势。

Comments Accepted by ICASSP2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12820 2026-01-21 cs.CV 81%

A Generalist Foundation Model for Total-body PET/CT Enables Diagnostic Reporting and System-wide Metabolic Profiling

一种通用的基础模型用于全身PET/CT,实现诊断报告和系统层面的代谢分析

Wei Chen, Liang Wu, Shuyi Lu, Yuanyuan Sun, Wenkai Bi, Zilong Yuan, Yaoyao He, Feng Wang, Junchi Ma, Shuyong Liu, Zhaoping Cheng, Xiaoyan Hu, Jianfeng Qiu

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);image-text(abstract);multimodal foundation model(abstract)

AI总结 SDF-HOLO是一种用于全身PET/CT的通用基础模型,通过多模态学习实现诊断报告和系统代谢分析,提升医疗影像处理的精度与效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.20136 2026-01-21 cs.CL cs.AI cs.IR 81%

Multi-Stage Verification-Centric Framework for Mitigating Hallucination in Multi-Modal RAG

多阶段验证导向框架用于缓解多模态RAG中的幻觉

Baiyu Chen, Wilson Wongso, Xiaoqian Hu, Yue Tan, Flora Salim

机构 * The University of New South Wales(新南威尔士大学)

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CL、cs.AI

AI总结 本文提出了一种多阶段验证导向框架,通过优先考虑事实准确性和真实性来缓解多模态RAG中的幻觉问题,并在KDD Cup 2025中取得第三名。

Comments KDD Cup 2025 Meta CRAG-MM Challenge: Third Prize in the Single-Source Augmentation Task

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.24816 2026-01-21 cs.CV cs.AI 81%

Perception, Understanding and Reasoning, A Multimodal Benchmark for Video Fake News Detection

感知、理解和推理,一个用于视频虚假新闻检测的多模态基准

Cui Yakun, Peng Qi, Fushuo Huo, Hang Du, Weijie Shi, Juntao Dai, Zhenghao Zhu, Sirui Han, Yike Guo

机构 * The Hong Kong University of Science and Technology(香港科学与技术大学) National University of Singapore(新加坡国立大学) The Hong Kong Polytechnic University(香港理工大学) Beijing University of Posts and Telecommunications(北京邮电大学) Peking University(北京大学)

专题命中 多模态评测 :multimodal(title);multi-modal(abstract);分类 cs.CV、cs.AI

AI总结 本文提出POVFNDB多模态基准,通过10个任务系统评估MLLMs在视频虚假新闻检测中的感知、理解和推理能力,并通过微调Qwen2.5VL-7B-Instruct达到最先进的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13453 2026-01-21 cs.CL cs.HC 79%

PhysicsSolutionAgent: Towards Multimodal Explanations for Numerical Physics Problem Solving

PhysicsSolutionAgent: 向数值物理问题求解的多模态解释迈进

Aditya Thole, Anmol Agrawal, Arnav Ramamoorthy, Dhruv Kumar

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CL

AI总结 PhysicsSolutionAgent通过生成多模态视频解释,提升数值物理问题的可视化教学效果,揭示了多模态推理与评估的挑战。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13052 2026-01-21 cs.CV 79%

GridNet-HD: A High-Resolution Multi-Modal Dataset for LiDAR-Image Fusion on Power Line Infrastructure

GridNet-HD: 一种高分辨率多模态数据集,用于电力线路基础设施的LiDAR-图像融合

Antoine Carreaud, Shanci Li, Malo De Lacour, Digre Frinde, Jan Skaloud, Adrien Gressin

专题命中 多模态评测 :multi-modal(title,abstract);分类 cs.CV

AI总结 GridNet-HD是一个高分辨率多模态数据集,用于电力线路基础设施的LiDAR-图像融合,通过融合高密度LiDAR和高分辨率影像提升3D语义分割性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.10949 2026-01-21 cs.CV 79%

MMedExpert-R1: Strengthening Multimodal Medical Reasoning via Domain-Specific Adaptation and Clinical Guideline Reinforcement

MMedExpert-R1: 通过领域特定适应与临床指南强化多模态医学推理

Meidan Ding, Jipeng Zhang, Wenxuan Wang, Haiqin Zhong, Xiaoling Luo, Wenting Chen, Linlin Shen

机构 * College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院) School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院) Guangdong Provincial Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室) The Hong Kong University of Science and Technology(香港科学与技术大学) Renmin University of China(中国人民大学) School of Biomedical Engineering, Shenzhen University(深圳大学生物医学工程学院)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 MMedExpert-R1通过领域特定适应和临床指南强化,提升多模态医学推理能力,实现多专科对齐和高精度推理性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.10360 2026-01-21 cs.CV 79%

FaceXBench: Evaluating Multimodal LLMs on Face Understanding

FaceXBench: 评估多模态大语言模型在面部理解上的能力

Kartik Narayan, Vibashan VS, Vishal M. Patel

机构 * Department of Computer Science, Johns Hopkins University(计算机科学系,约翰霍普金斯大学) Department of Electrical and Computer Engineering, Johns Hopkins University(电气与计算机工程系,约翰霍普金斯大学)

专题命中 多模态评测 :multimodal(title,abstract);分类 cs.CV

AI总结 FaceXBench通过5000个多模态问题评估MLLMs在面部理解上的能力,揭示了现有模型在复杂任务中的不足。

Comments Accepted in IEEE T-BIOM. Project Page: https://kartik-3004.github.io/facexbench/

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13639 2026-01-21 cs.RO 78%

A General One-Shot Multimodal Active Perception Framework for Robotic Manipulation: Learning to Predict Optimal Viewpoint

一种通用的一次性多模态主动感知框架用于机器人操作:学习预测最佳视角

Deyun Qin, Zezhi Liu, Hanqian Luo, Xiao Liang, Yongchun Fang

机构 * Institute of Robotics and Automatic Information Systems, College of Artificial Intelligence, Nankai University(机器人与自动信息系统研究所,人工智能学院,南开大学) Tianjin Key Laboratory of Intelligent Robotics, Nankai University(智能机器人重点实验室,南开大学) College of Artificial Intelligence, Nankai University(人工智能学院,南开大学) Department of Computing, The Hong Kong Polytechnic University(computing 部门,香港理工大学)

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本文提出了一种通用的一次性多模态主动感知框架,用于机器人操作,通过多模态特征融合和交叉注意力机制提升抓取成功率,并实现仿真到现实的无缝迁移。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.09853 2026-01-21 cs.LG 78%

ConSurv: Multimodal Continual Learning for Survival Analysis

ConSurv:面向生存分析的多模态持续学习

Dianzhi Yu, Conghao Xiong, Yankai Chen, Wenqian Cui, Xinni Zhang, Yifei Zhang, Hao Chen, Joseph J. Y. Sung, Irwin King

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 ConSurv是首个用于生存分析的多模态持续学习方法,通过多阶段混合专家和特征受限重放技术,有效解决灾难性遗忘和复杂模态交互问题。

Comments 14 pages, 4 figures. This is the extended version of the paper accepted at AAAI 2026, which includes all technical appendices and additional experimental details

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12584 2026-01-21 cond-mat.mtrl-sci 78%

Multi-modal data-driven microstructure characterization

多模态数据驱动的微观结构表征

Qi Zhang, Santiago Benito, Sebastian Weber, Markus Stricker

专题命中 多模态评测 :multi-modal(title);multimodal(abstract)

AI总结 本研究通过多模态数据驱动方法实现自动化的微观结构表征,包括晶粒分割和晶界检测,利用信息论优化参数选择和潜在空间特征映射。

Comments 23 pages, 14 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.11606 2026-01-21 cs.LG 78%

A Multimodal Data Processing Pipeline for MIMIC-IV Dataset

用于MIMIC-IV数据集的多模态数据处理流程

Farzana Islam Adiba, Varsha Danduri, Fahmida Liza Piya, Ali Abbasi, Mehak Gupta, Rahmatollah Beheshti

机构 * University of Delaware(德克萨斯大学)

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 本文提出了一种用于MIMIC-IV数据集的多模态数据处理流程,旨在提升多模态数据处理效率和研究可重复性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2509.23627 2026-01-21 cs.SI 78%

MASH: A Multiplatform and Multimodal Annotated Dataset for Societal Impact of Hurricane

MASH: 一种用于飓风社会影响的多平台和多模态注释数据集

Ruichen Yao, Aslanbek Murzakhmetov, Raaghav Pillai, Aliya Maussymbayeva, Zelin Li, Yifan Liu, Yaokun Liu, Lanyu Shang, Yang Zhang, Na Wei, Ximing Cai, Dong Wang

专题命中 多模态评测 :multimodal(title,abstract)

AI总结 MASH数据集通过多平台和多模态注释,全面覆盖飓风社会影响的多维度研究需求。

Comments preprint under review

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.14055 2026-01-21 cs.CV cs.AI 76%

Decoder-Free Supervoxel GNN for Accurate Brain-Tumor Localization in Multi-Modal MRI

无解码器的监督超像素图神经网络用于多模态MRI中脑肿瘤的准确定位

Andrea Protani, Marc Molina Van Den Bosch, Lorenzo Giusti, Heloisa Barbosa Da Silva, Paolo Cacace, Albert Sund Aillet, Miguel Angel Gonzalez Ballester, Friedhelm Hummel, Luigi Serio

机构 * European Organization for Nuclear Research, Geneva, Switzerland(欧洲核子研究中心) Dept. of Neuroscience, École Polytechnique Fédérale de Lausanne, Lausanne, Switzerland(瑞士洛桑联邦理工学院神经科学系) BCN Medtech, Dept. of Engineering, Universitat Pompeu Fabra, Barcelona, Spain(巴塞罗那大学工程系) Universidade de Coimbra, Coimbra, Portugal(科英布拉大学) Sapienza Università di Roma, Rome, Italy(罗马萨皮恩扎大学) ICREA, Barcelona, Spain(巴塞罗那高等科研机构)

专题命中 多模态评测 :multi-modal(title);分类 cs.CV、cs.AI

AI总结 本文提出SVGFormer,一种无解码器的监督超像素图神经网络,用于多模态MRI中脑肿瘤的准确定位,通过层次编码器结合Transformer和图注意力网络,实现高精度和可解释的医学图像分析。

Comments 10 pages, 3 figures,

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.06243 2026-01-21 cs.CL cs.AI 74%

CoT Referring: Improving Referring Expression Tasks with Grounded Reasoning

CoT Referring: 通过 grounded 推理改进指称表达任务

Qihua Dong, Luis Figueroa, Handong Zhao, Kushal Kafle, Jason Kuen, Zhihong Ding, Scott Cohen, Yun Fu

机构 * Adobe Research(Adobe研究院) Northeastern University(东北大学)

专题命中 多模态评测 :MLLM(abstract,comments);multimodal(abstract);分类 cs.CL、cs.AI

AI总结 通过 grounded 推理改进指称表达任务,提出CoT Referring方法,提升多模态大语言模型在复杂指称场景中的性能。

Comments MLLM, Referring Expression Segmentation

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.24866 2026-01-21 cs.CV 74%

TalkingHeadBench: A Multi-Modal Benchmark & Analysis of Talking-Head DeepFake Detection

TalkingHeadBench: 一个多模态基准及谈话头深度伪造检测分析

Xinqi Xiong, Prakrut Patel, Qingyuan Fan, Amisha Wadhwa, Sarathy Selvam, Xiao Guo, Luchao Qi, Xiaoming Liu, Roni Sengupta

机构 * University of North Carolina at Chapel Hill(北卡罗来纳大学教堂山分校) Michigan State University(密歇根州立大学)

专题命中 多模态评测 :multi-modal(title);分类 cs.CV

AI总结 TalkingHeadBench是一个多模态基准,用于评估和分析最先进的谈话头深度伪造检测方法,通过精心设计的协议和数据集提升检测模型的鲁棒性和泛化能力。

Comments WACV2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12256 2026-01-21 cs.AI 74%

Improving Large Molecular Language Model via Relation-aware Multimodal Collaboration

通过关系感知的多模态协作改进大型分子语言模型

Jinyoung Park, Minseong Bae, Jeehye Na, Hyunwoo J. Kim

机构 * Korea University(韩国大学)

专题命中 多模态评测 :multimodal(title);分类 cs.AI

AI总结 CoLLaMo通过关系感知的多模态协作机制提升分子语言模型的泛化能力,优化分子模态整合并改进评估方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.13299 2026-01-21 cs.CV 70%

Enginuity: Building an Open Multi-Domain Dataset of Complex Engineering Diagrams

Enginuity:构建一个复杂的工程图多领域开放数据集

Ethan Seefried, Prahitha Movva, Naga Harshita Marupaka, Tilak Kasturi, Tirthankar Ghosal

机构 * Oak Ridge National Laboratory(奥克伍德国家实验室)

专题命中 多模态评测 :multimodal(abstract);cross-modal(abstract);分类 cs.CV

AI总结 Enginuity是一个开放的多领域工程图数据集,旨在通过结构注释帮助AI处理工程图解析和科学发现。

Comments Accepted at the 39th Conference on Neural Information Processing Systems (NeurIPS 2025) Workshop: Ai4 Science

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12585 2026-01-21 cs.HC cs.AI cs.ET 70%

Do MLLMs See What We See? Analyzing Visualization Literacy Barriers in AI Systems

MLLMs是否看到我们看到的?分析AI系统中可视化素养障碍

Mengli, Duan, Yuhe, Jiang, Matthew Varona, Carolina Nobre

机构 * University of Toronto(多伦多大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.AI

AI总结 本研究分析了MLLMs在可视化素养中的障碍,揭示了两种机器特定的障碍,并展示了模型在复杂可视化任务上的表现差异。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.12505 2026-01-21 cs.CL 70%

DoPE: Decoy Oriented Perturbation Encapsulation Human-Readable, AI-Hostile Documents for Academic Integrity

DoPE: 伪装导向扰动封装用于学术诚信的人可读AI敌对文档

Ashish Raj Shekhar, Shiven Agarwal, Priyanuj Bordoloi, Yash Shah, Tejas Anvekar, Vivek Gupta

机构 * Arizona State University(亚利桑那州立大学)

专题命中 多模态评测 :multimodal(abstract);MLLM(abstract);分类 cs.CL

AI总结 DoPE通过在考试文档中嵌入语义伪装,利用MLLM pipelines的渲染-解析差异,实现对AI自动解决的预防和检测,提升学术诚信保障。

详情

展开后加载摘要…

URL PDF HTML 收藏