Aligning Visual Regions and Textual Concepts for Semantic-Grounded Image Representations
专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2019
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV
Comments Accepted by NeurIPS 2019
专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV
Comments Published at ICCV'2019
Journal ref The IEEE International Conference on Computer Vision (ICCV) 2019
专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV
Comments EMNLP 2019
专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV
Comments ICCV 2019 accepted paper
专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV
Comments 11 pages, 7 figures
专题命中 视觉问答 :visual reasoning(abstract);visual question answering(abstract);分类 cs.CV
专题命中 视觉问答 :visual reasoning(abstract);visual question answering(abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(abstract);grounding(abstract);分类 cs.CV
FashionKG-RAG:面向时尚问答的知识图谱增强检索增强生成
专题命中 视觉问答 :grounding(abstract,abstract_cn)
AI总结 针对现有时尚知识图谱的局限性,本文提出全领域知识图谱FashionEcoKG,并开发无需训练的PG-RAG框架,通过双粒度路径重排序模块提升时尚问答的检索与答案准确性,效果优于多种基线方法。
答案保留型攻击下的模型置信度:信息性-可操纵性前沿
专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract)
AI总结 该研究针对答案保留型攻击,发现视觉-语言系统的置信度信号不具备固有鲁棒性,四类防御均无效,协调攻击可大幅降低置信度门控下的接受准确率。
基于大语言模型的视觉编码器分层预训练
机构 * University of Cincinnati(辛辛那提大学) ; National Yang Ming Chiao Tung University(国立阳明交通大学)
专题命中 视觉问答 :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 本文提出HIVE框架,通过引入视觉编码器与大语言模型间的分层交叉注意力机制,提升视觉语言对齐,改进特征融合与表征学习,实验表明其在图像分类和多模态任务中表现优异。
Comments 17 pages, 14 figures, accepted to Computer Vision and Pattern Recognition Conference (CVPR) Workshops 2026. 5th MMFM Workshop: What is Next in Multimodal Foundation Models?
Journal ref In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (pp. 7415-7424) 2026
K12-KGraph:一种对齐课程的图谱用于基准测试和训练教育大语言模型
机构 * Peking University(北京大学) ; Institute for Advanced Algorithms Research(先进算法研究所) ; OriginHub Technology(OriginHub技术) ; Zhongguancun Academy(中关村学院)
专题命中 视觉问答 :vision-language model(abstract);grounding(abstract)
AI总结 本文提出K12-KGraph,基于中小学教材构建的课程对齐知识图谱,用于构建多选基准测试K12-Bench和训练数据集K12-Train,验证课程结构监督在教育模型训练中的高效性。
不确定性并非临床VQA的安全网,但它能预测模型失败吗?
机构 * Amsterdam University Medical Center, University of Amsterdam(阿姆斯特丹大学医学中心) ; Amsterdam Public Health(阿姆斯特丹公共卫生) ; LMU Munich(慕尼黑大学) ; Munich Center for Machine Learning (MCML)(慕尼黑机器学习中心)
专题命中 视觉问答 :vision-language model(abstract);VLM(abstract_cn)
AI总结 研究临床视觉语言模型的不确定性估计是否可靠,发现其质量随模型准确率变化,在模型脆弱时失效,但能预测扰动下的性能崩溃。
Comments 17 pages, 4 figures
将具身问答从感知扩展到决策
机构 * Peking University(北京大学) ; XYZ Embodied AI(XYZ具身AI)
专题命中 视觉问答 :VLM(abstract,abstract_cn)
AI总结 提出大规模具身问答数据集EQA-Decision和基线模型RoboDecision,系统覆盖静态场景构建、空间理解、任务动态推理和即时决策四个维度,以统一框架评估具身环境中的感知、推理和行动级决策。
Comments 11 pages,4 figures
一种用于柬埔寨检索增强问答的语言模型比较研究
机构 * Department of Computer Science, Chungbuk National University(Chungbuk National University 计算机科学系) ; Department of Big Data, Chungbuk National University(Chungbuk National University 大数据系) ; General Department of Information and Communication Technology, Ministry of Post and Telecommunications(邮电部信息和通信技术总局) ; Department of Management Information Systems, Chungbuk National University(Chungbuk National University 管理信息系统系) ; BigDatalabs Co., Ltd(BigDatalabs 公司)
专题命中 视觉问答 :grounding(abstract,abstract_cn)
AI总结 本文针对低资源非拉丁语种柬埔寨语言,比较了多种语言模型在检索增强问答任务中的性能,发现检索器选择是影响效果的关键因素,生成器在不同指标上表现各异。
Comments 14 pages, 1 figure,
大语言模型的瓶颈:为何开源视觉大语言模型在层级视觉识别上遇到困难
机构 * Boston University(波士顿大学)
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 本文指出开源大语言模型缺乏对视觉世界的层级知识,导致视觉大语言模型在识别如水母鱼但无法识别脊椎动物时存在瓶颈,通过构建六种分类学和四个图像数据集的百万级多项选择视觉问答任务验证了这一问题。
Comments Accepted to CVPR 2026. Project page and code: https://yuanqing-ai.github.io/llm-hierarchy/
SEA-Vision:面向东南亚的多语言综合文档和场景文本理解基准
机构 * Xiamen University, China(厦门大学,中国) ; Shopee, China(Shopee,中国) ; Tongji University, China(同济大学,中国)
专题命中 视觉问答 :visual question answering(abstract);MLLM(abstract)
AI总结 SEA-Vision提出一个多语言基准,用于评估文档解析和文本中心视觉问答,涵盖11种东南亚语言,包含15234页文档和7496对问答对,揭示多语言文档理解的差距。
Comments Accepted By CVPR2026
自然语言基础的思维社会中的风暴
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 本文探讨了基于自然语言的思维社会(NLSOMs)的结构和应用,通过多代理系统解决多种AI任务,并提出未来研究方向。
Comments published in Computational Visual Media Journal (CVMJ); 9 pages in main text + 7 pages of references + 38 pages of appendices, 14 figures in main text + 13 in appendices, 7 tables in appendices
Journal ref (2025). Computational Visual Media, 11(1), 29-81
ExStrucTiny:一种用于从文档图像中进行模式变量结构信息提取的基准
机构 * J.P. Morgan AI Research(摩根大通人工智能研究)
专题命中 视觉问答 :vision language model(abstract);visual question answering(abstract)
AI总结 ExStrucTiny是一个新的文档图像结构信息提取基准,通过结合手动和合成样本,涵盖更多多样化的文档类型和提取场景,旨在提升通用模型在结构化信息提取中的性能。
Comments EACL 2026, main conference
TabRAG:通过结构化表示改进表格文档问答以增强检索增强生成
专题命中 视觉问答 :vision language model(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 TabRAG通过结构化表示改进表格文档问答,采用布局分割和视觉语言模型解析,提升表格问答性能。
Comments NeurIPS 2025 AI4Tab
DoPE: 伪装导向扰动封装用于学术诚信的人可读AI敌对文档
机构 * Arizona State University(亚利桑那州立大学)
专题命中 视觉问答 :multimodal large language model(abstract);MLLM(abstract)
AI总结 DoPE通过在考试文档中嵌入语义伪装,利用MLLM pipelines的渲染-解析差异,实现对AI自动解决的预防和检测,提升学术诚信保障。
集成多个基于VLLM内部表示的幻觉检测器
专题命中 视觉问答 :VLM(abstract);visual question answering(abstract)
AI总结 本文提出通过集成多个基于VLLM内部表示的幻觉检测模型,以减少幻觉并提高VQA任务的准确性。
Comments 5th place solution at Meta KDD Cup 2025
LinkedOut: 从视频大语言模型中提取世界知识表示以实现下一代视频推荐
机构 * Northeastern University(东北大学) ; LinkedIn(领英) ; University of Wisconsin–Madison(威斯康星大学麦迪逊分校)
专题命中 视觉问答 :visual reasoning(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 LinkedOut通过从视频中提取世界知识表示,实现低延迟、多视频输入的视频推荐,无需人工标注,取得最佳性能。
先思后定(ThiFAN-VQA):一种用于灾后损害评估的两阶段链式思考框架
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
AI总结 ThiFAN-VQA通过两阶段推理框架提升灾后损害评估的准确性与可解释性。
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted at AAAI'26
专题命中 视觉问答 :VLM(abstract);visual question answering(abstract)
机构 * China University of Petroleum-Beijing(中国石油大学(北京)) ; Leyard Optoelectronic(莱亚德光电) ; University of Wisconsin-Milwaukee(威斯康星大学密尔沃基分校)
专题命中 视觉问答 :visual question answering(abstract);分类 cs.CV、cs.AI、cs.LG
Comments Accepted by ICCV 2025
专题命中 视觉问答 :multimodal large language model(abstract);MLLM(abstract)
Comments 6 pages, 3 figures
机构 * Tianjin University(天津大学) ; Shandong Institute of Petroleum and Chemical Technology(山东石油化学工业技术研究所) ; Beijing Institute of Technology(北京理工大学)
专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract)
Comments I need to modify the content of the article