MemoryCard: Topic-Aware Multi-Modal Clue Compression for Long-Video Question Answering
MemoryCard: 面向长视频问答的主题感知多模态线索压缩
Qing Yang, Pengcheng Huang, Xinze Li, Zhenghao Liu, Yukun Yan, Yu Gu, Ge Yu, Gang Li, Maosong Sun
机构
*
School of Computer Science and Engineering, Northeastern University(东北大学计算机科学与工程学院)
;
Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系)
;
Digital China Group(数字中国集团)
Cognitive Chain-of-Thought (CoCoT): Structured Multimodal Reasoning about Social Situations
认知链式推理(CoCoT):关于社会情境的结构化多模态推理
Eunkyu Park, Wesley Hanwen Deng, Gunhee Kim, Motahhare Eslami, Maarten Sap
机构
*
Seoul National University(首尔国立大学)
;
Human-Computer Interaction Institute, Carnegie Mellon University(人机交互研究所,卡内基梅隆大学)
;
Language Technologies Institute, Carnegie Mellon University(语言技术研究所,卡内基梅隆大学)
Beyond Confidence: Test-Time Scaling for Multi-Turn Search Agents via Retrieval Grounding
超越置信度:基于检索 grounding 的多轮搜索智能体测试时缩放
Hyunho Kook, Junhyuk So, Tianyu Fu, Haizhong Zheng, Beidi Chen
机构
*
University of Southern California(南加州大学)
;
Pohang University of Science and Technology (POSTECH)(浦项科技大学)
;
Tsinghua University(清华大学)
;
Carnegie Mellon University(卡内基梅隆大学)
Overview of SHROOM-Visions 2026: A Shared Task on Hallucination Detection in Large Vision-Language Models
SHROOM-Visions 2026概览:大型视觉语言模型幻觉检测共享任务
Raúl Vázquez, Aman Sinha, Chuyuan Li, Artem Shelmanov, Artem Vazhentsev, Claudio Savelli, Eduardo Calò, Emilio Raimond, Stella Frank, Hengyu Luo, Flavio Giobergia, Vincent Segonne, Lorenzo Vaiani, Jörg Tiedemann, Timothee Mickus
机构
*
University of Helsinki(赫尔辛基大学)
;
Politecnico di Torino(都灵理工大学)
;
Université Bretagne Sud(南布列塔尼大学)
;
University of Copenhagen(哥本哈根大学)
;
University Grenoble Alpes(格勒诺布尔阿尔卑斯大学)
;
University of Lorraine(洛林大学)
GRAFT: Grounded and Efficient Online Reinforcement Adaptation for Fine-Grained Robot Manipulation
GRAFT:面向精细机器人操作的 grounded 高效在线强化适应
Yibo Qiu, Haoliang Ye, Shu'ang Sun, Zan Huang, Ronald X Xu, Mingzhai Sun
机构
*
Suzhou Institute for Advanced Research, University of Science and Technology of China(中国科学技术大学苏州高等研究院)
;
School of Biomedical Engineering, Division of Life Sciences and Medicine, University of Science and Technology of China(中国科学技术大学生命科学与医学部生物医学工程学院)
Action- and Language-Conditioned Video Assessment for Embodied Control
面向具身控制的动作与语言条件视频评估
Hwanhee Kim, Jaehyun Jang, Seungmin Cha, Hyeonseo Yun, Donghoon Lee, Chang D. Yoo
机构
*
Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
;
School of Electrical Engineering, Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院电气工程学院)
机构
*
University of Notre Dame(notre dame 大学)
;
University of Washington(华盛顿大学)
;
Johns Hopkins University(约翰霍普金斯大学)
;
Georgia Institute of Technology(佐治亚理工学院)
专题命中
幻觉与鲁棒性
:multimodal large language model(title,abstract);MLLM(abstract,abstract_cn);分类 cs.AI
Comments17 pages (9 pages of main text, plus references and appendix), 7 figures, and 13 tables. Accepted to EMNLP 2026 Main Conference. Project page: this https URL (https://lizesheng13.github.io/Grasp-VL/)
PRISM: Self-Pruning Intrinsic Selection Method for Training-Free Multimodal Data Selection
PRISM:免训练多模态数据选择的自剪枝内在选择方法
Jinhe Bi, Aniri, Zengjie Jin, Yifan Wang, Danqi Yan, Wenke Huang, Xiaowen Ma, Sikuan Yan, Artur Hecker, Mang Ye, Xun Xiao, Hinrich Schuetze, Volker Tresp, Yunpu Ma
机构
*
LMU Munich(慕尼黑大学)
;
Munich Research Center, Huawei Technologies(慕尼黑研究中心,华为技术)
;
METEOR
;
School of Computer Science, Wuhan University(武汉大学计算机学院)
;
Munich Center for Machine Learning(慕尼黑机器学习中心)
专题命中
VLM训练与架构
:multimodal large language model(abstract);分类 cs.CV、cs.AI