Finding the Correct Visual Evidence Without Forgetting: Mitigating Hallucination in LVLMs via Inter-Layer Visual Attention Discrepancy
在不遗忘的情况下寻找正确的视觉证据:通过层间视觉注意力差异减轻LVLMs中的幻觉
Yutong Xie, Zhenglin Hua, Ran Wang, Wing W. Y. Ng, Xizhao Wang, Yuheng Jia
机构
*
School of Computer Science and Engineering, Southeast University, Nanjing, China(东南大学计算机科学与工程学院)
;
School of Artificial Intelligence, Shenzhen University, Shenzhen, China(深圳大学人工智能学院)
;
College of Computer Science and Software Engineering, Shenzhen University, Shenzhen, China(深圳大学计算机科学与软件工程学院)
;
Engineering, South China University of Technology, Guangzhou, China(华南理工大学工程学院)
;
National Engineering Laboratory for Big Data Systems Computing Technology, Shenzhen University, Shenzhen, China(深圳大学大数据系统计算技术国家工程实验室)
;
Key Laboratory of New Generation Artificial Intelligence Technology and Its Interdisciplinary Applications (Southeast University), Ministry of Education, China(新一代人工智能技术及其交叉应用重点实验室(东南大学),中华人民共和国教育部)
Epistemic Uncertainty Quantification for Pre-trained VLMs via Riemannian Flow Matching
通过黎曼流匹配对预训练视觉语言模型进行知识不确定性量化
Li Ju, Mayank Nautiyal, Andreas Hellander, Ekta Vats, Prashant Singh
机构
*
Department of Information Technology, Uppsala University, Uppsala, Sweden(瑞典乌普萨拉大学信息科技系)
;
Science for Life Laboratory, Uppsala University, Uppsala, Sweden(瑞典乌普萨拉大学生命科学实验室)
机构
*
Australian Institute for Machine Learning(澳大利亚机器学习研究院)
;
Responsible AI Research Centre(负责任人工智能研究中心)
;
College of Computer Science and Artificial Intelligence(计算机科学与人工智能学院)
;
Department of Data Science and AI(数据科学与人工智能部门)
;
School of Computer Science and Engineering(计算机科学与工程学院)
机构
*
State Key Lab. of AI Safety, Institute of Computing Technology, Chinese Academy of Sciences(中国科学院人工智能安全国家重点实验室,计算技术研究所)
;
School of Advanced Interdisciplinary Sciences, University of Chinese Academy of Sciences(中国科学院大学先进交叉科学学院)
;
School of Computer Science and Technology, University of Chinese Academy of Sciences(中国科学院大学计算机科学与技术学院)
专题命中
知识编辑与模型理解
:language model(abstract);分类 cs.LG
AI总结
本文提出了一种名为Locate-Then-Sparsify for Feature Steering (LTS-FS)的框架,通过定位和稀疏化策略,根据每层与幻觉的相关性调整特征引导强度,从而有效缓解视觉语言模型中的幻觉问题,同时保持良好的性能。
Revealing the Gap in Human and VLM Scene Perception through Counterfactual Semantic Saliency
通过反事实语义显著性揭示人类与VLM场景感知之间的差距
Ziqi Wen, Parsa Madinei, Miguel P. Eckstein
机构
*
Department of Computer Science, University of California, Santa Barbara(加州大学圣巴巴拉分校计算机科学系)
;
Department of Psychological and Brain Sciences, University of California, Santa Barbara(加州大学圣巴巴拉分校心理学与脑科学系)
FORGE: Fragment-Oriented Ranking and Generation for Context-Aware Molecular Optimization
FORGE:面向上下文的分子优化片段排序与生成
Qingchuan Zhang, He Cao, Hao Li, Yanjun Shao, Zhiyuan Liu, Shihang Wang, Shufang Xie, Shenghua Gao, Xinwu Ye
机构
*
University of Science and Technology of China(中国科学技术大学)
;
International Digital Economy Academy(国际数字经济学院)
;
Peking University(北京大学)
;
Yale University(耶鲁大学)
;
National University of Singapore(新加坡国立大学)
;
Macao Polytechnic University(澳门理工学院)
;
Zhongguancun Academy(中关村学院)
;
University of Hong Kong(香港大学)
机构
*
Laboratory of Molecular Pharmacokinetics, Graduate School of Pharmaceutical Sciences, The University of Tokyo(分子药代动力学实验室,药学研究生院,东京大学)
;
The Institute of Statistical Mathematics (ISM), Research Organization of Information and Systems(统计数学研究所(ISM),信息与系统研究组织)
机构
*
Center for Applied Mathematics, Cornell University(康奈尔大学应用数学中心)
;
Department of Mathematics, Imperial College London(伦敦帝国理工学院数学系)
;
School of Electrical and Computer Engineering, Cornell University(康奈尔大学电气与计算机工程学院)
Toward Better Geometric Representations for Molecule Generative Models
迈向更优的分子生成模型几何表示
Shaoheng Yan, Zian Li, Cai Zhou, Qiaojing Huang, Kai Liu, Muhan Zhang
机构
*
Institute for Artificial Intelligence, PKU(北京大学人工智能学院)
;
ByteDance AI Drug Discovery(字节跳动AI药物发现)
;
School of Intelligence Science and Technology, PKU(北京大学智能科学与技术学院)
;
Anew Labs(Anew实验室)
;
Department of Electrical Engineering and Computer Science, MIT(麻省理工学院电子工程与计算机科学系)
;
State Key Laboratory of General Artificial Intelligence, PKU(北京大学通用人工智能国家重点实验室)
Neurosymbolic Framework for Concept-Driven Logical Reasoning in Skeleton-Based Human Action Recognition
基于骨架的人体动作识别的概念驱动逻辑推理神经符号框架
Talha Ilyas, Deval Mehta, Zongyuan Ge
机构
*
Department of ECSE, Faculty of Engineering, Monash University, Australia(莫纳什大学工程学院电子与计算机工程系,澳大利亚)
;
AIM for Health Lab, Faculty of Information Technology, Monash University, Australia(莫纳什大学信息科技学院健康人工智能实验室,澳大利亚)
;
Department of DSAI, Faculty of Information Technology, Monash University, Australia(莫纳什大学信息科技学院数据科学与人工智能系,澳大利亚)