MM-SCALE: Grounded Multimodal Moral Reasoning via Scalar Judgment and Listwise Alignment
MM-SCALE: 基于标量判断和列表对齐的 grounded 多模态道德推理
Eunkyu Park, Wesley Hanwen Deng, Cheyon Jin, Matheus Kunzler Maldaner, Jordan Wheeler, Jason I. Hong, Hong Shen, Adam Perer, Ken Holstein, Motahhare Eslami, Gunhee Kim
机构
*
Seoul National University(首尔国立大学)
;
Carnegie Mellon University(卡内基梅隆大学)
;
University of Florida(佛罗里达大学)
;
Epic Games
Training-Free and Interpretable Hateful Video Detection via Multi-stage Adversarial Reasoning
无需训练的多阶段对抗推理 hateful 视频检测
Shuonan Yang, Yuchen Zhang, Zeyu Fu
机构
*
Multimodal Intelligence Lab, Department of Computer Science, University of Exeter, United Kingdom(埃克塞特大学计算机科学系多模态智能实验室)
;
Institute for Analytics and Data Science, University of Essex, United Kingdom(埃塞克斯大学分析与数据科学研究所)
专题命中
推理评测
:reasoning(title,abstract)
AI总结
MARS通过多阶段对抗推理框架实现无需训练的可解释仇恨视频检测,提升检测可靠性与透明度。
CommentsAccepted at ICASSP 2026. \c{opyright} 2026 IEEE. This is the author accepted manuscript. The final published version will be available via IEEE Xplore
机构
*
College of Computer Science and Software Engineering, Shenzhen University(深圳大学计算机科学与软件工程学院)
;
School of Artificial Intelligence, Shenzhen University(深圳大学人工智能学院)
;
Guangdong Provincial Key Laboratory of Intelligent Information Processing(广东省智能信息处理重点实验室)
;
The Hong Kong University of Science and Technology(香港科学与技术大学)
;
Renmin University of China(中国人民大学)
;
School of Biomedical Engineering, Shenzhen University(深圳大学生物医学工程学院)
VisualQuest: A Benchmark for Abstract Visual Reasoning in MLLMs
VisualQuest: 一个用于多模态大语言模型(MLLMs)抽象视觉推理的基准数据集
Kelaiti Xiao, Liang Yang, Dongyu Zhang, Paerhati Tulajiang, Hongfei Lin
机构
*
School of Computer Science and Technology, Dalian University of Technology, Dalian, China(大连理工大学计算机科学与技术学院)
;
School of Foreign Languages, Dalian University of Technology, Dalian, China(大连理工大学外语学院)
;
School of Computer Science and Technology, Xinjiang Normal University, Urumqi, China(新疆师范大学计算机科学与技术学院)
Synthetic Vasculature and Pathology Enhance Vision-Language Model Reasoning
合成血管和病理增强视觉-语言模型推理
Chenjun Li, Cheng Wan, Laurin Lux, Alexander Berger, Richard B. Rosen, Martin J. Menten, Johannes C. Paetzold
机构
*
Cornell University(康奈尔大学)
;
Weill Cornell Medicine(韦尔·康奈尔医学)
;
Technical University of Munich(慕尼黑技术大学)
;
New York Eye and Ear Infirmary of Mount Sinai(圣文森特医院)
;
Cornell Tech(康奈尔科技)
Video-CoM: Interactive Video Reasoning via Chain of Manipulations
Video-CoM:通过操作链进行交互式视频推理
Hanoona Rasheed, Mohammed Zumri, Muhammad Maaz, Ming-Hsuan Yang, Fahad Shahbaz Khan, Salman Khan
机构
*
Mohamed bin Zayed University of AI(穆罕默德·本·扎耶德人工智能大学)
;
University of California Merced(加州梅尔 Ced 大学)
;
Google Research(谷歌研究院)
;
Linköping University(林奈大学)
;
Australian National University(澳大利亚国立大学)
机构
*
MoE Key Laboratory of Brain-inspired Intelligent Perception and Cognition, University of Science and Technology of China(脑启发智能感知与认知国家重点实验室,中国科学技术大学)
;
State Key Laboratory of Multimedia Information Processing, School of Computer Science, Peking University(多媒体信息处理国家重点实验室,北京大学计算机学院)
;
CUHK(香港大学)