机构
*
State Key Laboratory of Cognitive Intelligence(认知智能国家重点实验室)
;
University of Science and Technology of China(中国科学技术大学)
;
Artificial Intelligence Research Institute(人工智能研究院)
;
iFLYTEK Co., Ltd(iFLYTEK公司)
专题命中
视觉问答
:multimodal large language model(abstract);分类 cs.AI
CommentsThis paper was accepted and scheduled for inclusion in the ICALT 2025 proceedings but was ultimately not published due to absence from the conference presentation. It appears in the official program booklet. Conference: 2025 IEEE International Conference on Advanced Learning Technologies (ICALT)
Enhancing Radiology Report Generation and Visual Grounding using Reinforcement Learning
通过强化学习增强放射学报告生成和视觉基础识别
Benjamin Gundersen, Nicolas Deperrois, Samuel Ruiperez-Campillo, Thomas M. Sutter, Julia E. Vogt, Michael Moor, Farhad Nooralahzadeh, Michael Krauthammer
机构
*
University of Zurich(苏黎世大学)
;
ETH Zurich(苏黎世联邦理工学院)
;
Zurich University of Applied Sciences(苏黎世应用科学大学)
AutoMedic: An Automated Evaluation Framework for Clinical Conversational Agents with Medical Dataset Grounding
AutoMedic: 一种用于具有医学数据集支撑的临床对话代理的自动化评估框架
Gyutaek Oh, Sangjoon Park, Byung-Hoon Kim
机构
*
Yonsei University College of Medicine(延世大学医学院)
;
Yonsei Institute for Digital Health(延世大学数字健康研究所)
;
Yonsei University(延世大学)
;
Institute of Behavioral Sciences in Medicine(医学行为科学研究所)
机构
*
Shanghai AI Laboratory(上海人工智能实验室)
;
Shanghai Jiaotong University(上海交通大学)
;
The Chinese University of Hong Kong(香港中文大学)
;
Zhejiang University(浙江大学)
;
Beihang University(北京航空航天大学)
;
Xi’an Jiaotong University(西安交通大学)
;
University of Hong Kong(香港大学)
;
Fudan University(复旦大学)
;
University of California, Los Angeles(加州大学洛杉矶分校)
ConStruct: Structural Distillation of Foundation Models for Prototype-Based Weakly Supervised Histopathology Segmentation
ConStruct:用于基于原型的弱监督病理分割的基模态结构蒸馏
Khang Le, Ha Thach, Anh M. Vu, Trang T. K. Vo, Han H. Huynh, David Yang, Minh H. N. Le, Thanh-Huy Nguyen, Akash Awasthi, Chandra Mohan, Zhu Han, Hien Van Nguyen
机构
*
Ho Chi Minh City University of Technology(胡志明市技术大学)
;
University of Technology Sydney(悉尼技术大学)
;
University of Houston(休斯顿大学)
;
University of Information Technology(信息科技大学)
;
College of Medical Science and Technology, Taipei Medical University(台北医学院医学科技学院)
;
Department of Computer Science, Emory University(埃默里大学计算机科学系)
;
Montefiore Medical Center, Albert Einstein College of Medicine(蒙特福伊医疗中心,阿尔伯特·爱因斯坦医学院)
;
School of Computer Science, Carnegie Mellon University(卡内基梅隆大学计算机科学学院)
Comments15 pages, 6 figures, 1 table; accepted for AI-2025 Forty-fifth SGAI International Conference on Artificial Intelligence CAMBRIDGE, ENGLAND 16-18 DECEMBER 2025
SAVE: Sparse Autoencoder-Driven Visual Information Enhancement for Mitigating Object Hallucination
SAVE:基于稀疏自编码器的视觉信息增强用于缓解物体幻觉
Sangha Park, Seungryong Yoo, Jisoo Mok, Sungroh Yoon
机构
*
Department of Electrical and Computer Engineering, Seoul National University(电子与计算机工程系,首尔国立大学)
;
Daegu Gyeongbuk Institute of Science and Technology(大邱庆州科学技术院)
;
IPAI, AIIS, ASRI, INMC, and ISRC, Seoul National University(IPAI、AIIS、ASRI、INMC 和 ISRC,首尔国立大学)
专题命中
幻觉与鲁棒性
:multimodal large language model(abstract);分类 cs.CV、cs.AI
Beyond Pixels: A Training-Free, Text-to-Text Framework for Remote Sensing Image Retrieval
超越像素:一种无训练的文本到文本框架用于遥感图像检索
J. Xiao, Y. Guo, X. Zi, K. Thiyagarajan, C. Moreira, M. Prasad
机构
*
Information Technology University of Technology Sydney Sydney, Australia(信息科技技术大学悉尼分校悉尼澳大利亚)
;
Robotics Laboratory (SensR Lab) Centre for Advanced Manufacturing Technology Western Sydney University Sydney, Australia(机器人实验室(SensR实验室)先进制造技术中心西悉尼大学悉尼澳大利亚)
;
The Data Science Institute Faculty of Engineering(数据科学学院工程学院)
Exploring Automated Recognition of Instructional Activity and Discourse from Multimodal Classroom Data
探索多模态课堂数据中教学活动和话语的自动化识别
Ivo Bueno, Ruikun Hou, Babette Bühler, Tim Fütterer, James Drimalla, Jonathan Kyle Foster, Peter Youngs, Peter Gerjets, Ulrich Trautwein, Enkelejda Kasneci