Show, Ask, Attend, and Answer: A Strong Baseline For Visual Question Answering
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments 11 pages, 4 figures, accepted in CVPR 2017 (poster)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments 11 pages, 7 figures, 3 tables in 2016 Conference on Neural Information Processing Systems (NIPS)
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments 14 pages. arXiv admin note: text overlap with arXiv:1511.06973
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments Submitted to IbPRIA'17, 8 pages, 3 figures, 1 table
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments The first three authors contributed equally. International Conference on Computer Vision (ICCV) 2015
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments 25 pages
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments 5 pages, 4 figures, 3 tables, presented at 2016 ICML Workshop on Human Interpretability in Machine Learning (WHI 2016), New York, NY. arXiv admin note: substantial text overlap with arXiv:1606.03556
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments 9 pages, 6 figures, 3 tables; Under review at EMNLP 2016
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments Accepted to IEEE Conf. Computer Vision and Pattern Recognition
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments Submitted to CVPR2016
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments One comparison method's scores are put into the correct column, and a new experiment of generating attention map is added
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
专题命中 视觉问答 :visual question answering(title,abstract);分类 cs.CV
Comments 20 pages
ReXSonoVQA:面向以过程为中心的超声理解的视频问答基准
机构 * Harvard Medical School(哈佛医学院)
专题命中 视觉问答 :LLaVA(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 ReXSonoVQA是一个包含514个视频片段和514个问题的视频问答基准,旨在评估动作-目标推理、artifact解决与优化及过程上下文与规划能力,测试VLMs在动态过程理解中的表现。
Journal ref Proceedings of the 7th Conference on Health, Inference, and Learning, PMLR 333:427-447, 2026
用于CT扫描中可靠且可审计的空间关系验证的模块化智能体
机构 * Ulm University(乌尔姆大学) ; Ulm University Hospital(乌尔姆大学医院) ; Technical University of Munich (TUM)(慕尼黑工业大学(TUM)) ; TUM University Hospital(慕尼黑工业大学医院) ; Department of Radiation Oncology, TUM University Hospital(慕尼黑工业大学医院放射肿瘤科)
专题命中 视觉问答 :VLM(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 针对医学视觉-语言模型空间推理薄弱的问题,提出模块化医学影像智能体,通过分阶段处理实现CT扫描空间关系验证,性能优于端到端基线,可作为未来医学影像智能体的构建块。
Journal ref Published at the MICCAI 2026 Agentic AI for Medicine Workshop
RISE:跨三维跟踪与结构化视觉-语言推理的路边基础设施序列理解
专题命中 视觉问答 :MLLM(abstract,abstract_cn);grounding(abstract);分类 cs.CV、cs.AI
AI总结 本研究提出RISE框架,结合仅图像的三维跟踪方法与结构化视觉-语言推理,构建含33910个问答对的RISE-VQA数据集,通过RISE-Bench评估任务,揭示相关挑战并验证域自适应与时间上下文的益处。
视频-大语言模型(Video-LLM)问答中幻觉引用的检测:自验证流水线与验证器 ablation 研究
专题命中 视觉问答 :grounding(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 针对视频-LLM问答中引用幻觉问题,提出自验证流水线,用小型自然语言推理模型作稳定验证器,在对抗性错误前提问题上检测率达79%,并发布相关代码。
360度图像感知与MLLMs:一个全面的基准和无需训练的方法
机构 * GSIS, Tohoku University(东大GSIS研究所,东京东大大学) ; RIKEN AIP, Japan(日本RIKEN AIP)
专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 本文提出360Bench基准和Free360方法,评估MLLMs在360度图像感知中的能力,揭示其不足,并提出基于场景图的无需训练框架提升VQA性能。
Journal ref ECCV2026 (Link: https://tranhuyen1191.github.io/360Bench-Free360/)
拼接与讲述:一种结构化多模态数据增强方法用于空间理解
机构 * School of Computer Science, Beijing Institute of Technology(北京理工大学计算机科学学院) ; School of Software and Microelectronics, Peking University(北京大学软件与微电子学院) ; Xiaohongshu Inc(小红书公司)
专题命中 视觉问答 :LLaVA(abstract,abstract_cn);vision-language model(abstract);分类 cs.CV、cs.AI
AI总结 Stitch and Tell通过结构化空间监督提升视觉-语言模型的空间理解能力,有效缓解空间幻觉并提高相关任务性能。
Journal ref Advances in Neural Information Processing Systems 38 (NeurIPS 2025)
CARVE:用于高效3D医学体积理解的视觉证据跨切片各向异性重分配
专题命中 视觉问答 :MLLM(summary_cn,abstract_cn);分类 cs.CV、cs.AI
AI总结 针对3D医学体积理解中切片式MLLM的视觉令牌冗余问题,提出无需训练的CARVE框架,通过跨切片各向异性重分配压缩80%令牌,在AMOS-MM等基准上性能优于现有方法。
ClinFusion:用于整体医学理解的以视觉为中心的多模态大语言模型系统
机构 * DAMO Academy, Alibaba Group(达摩院,阿里巴巴集团) ; Hupan Laboratory(湖畔实验室) ; College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院) ; Department of Computer Science and Technology, Tsinghua University(清华大学计算机科学与技术系) ; Department of Radiology, The Affiliated Yangming Hospital of Ningbo University(宁波大学附属阳明医院放射科) ; Zhejiang University-University of Illinois Urbana-Champaign Institute, Zhejiang University(浙江大学伊利诺伊大学厄巴纳香槟校区联合学院,浙江大学) ; Hepato-Pancreato-Biliary Center, Beijing Tsinghua Changgung Hospital, School of Clinical Medicine, Tsinghua Medicine, Tsinghua University(清华长庚医院肝胆胰中心,清华大学临床医学院,清华医学,清华大学) ; School of Software, Tsinghua University(清华大学软件学院) ; Beijing National Research Center for Information Science and Technology, Tsinghua University(清华大学北京信息科学与技术国家研究中心)
专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV、cs.AI
AI总结 研究针对多模态大语言模型在医学领域部署的挑战,提出ClinFusion,采用组合级联视觉编码器架构和视觉基础评估框架,在多模态医学基准测试中表现优异,超越开源和专有模型,经专家盲评验证效果良好。
Comments Code: https://github.com/alibaba-damo-academy/ClinFusion Models: https://huggingface.co/collections/Alibaba-DAMO-Academy/clinfusion
PathScale-R1:用于病理图像分析的跨尺度推理
机构 * National University of Singapore(新加坡国立大学) ; PuzzleLogic Pte Ltd(拼图逻辑私人有限公司) ; Fujian Medical University Cancer Hospital & Fujian Cancer Hospital(福建医科大学附属肿瘤医院)
专题命中 视觉问答 :vision-language model(abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV、cs.AI
AI总结 研究针对病理诊断多尺度需求及现有模型局限,引入抗捷径跨尺度病理推理框架,设计相关策略构建PathScale-VQA基准,经优化得到PathScale-R1,实验证明其在跨尺度推理任务性能及对单尺度病理VQA的有效迁移。
增强病理视觉语言模型的跨尺度推理能力
机构 * Department of Electrical and Computer Engineering, National University of Singapore(新加坡国立大学电气与计算机工程系) ; PuzzleLogic Pte Ltd(PuzzleLogic私人有限公司) ; Department of Pathology, Fujian Medical University Cancer Hospital & Fujian Cancer Hospital(福建医科大学附属肿瘤医院病理科暨福建省肿瘤医院)
专题命中 视觉问答 :vision-language model(abstract);VLM(abstract_cn);visual question answering(abstract);分类 cs.CV、cs.AI
AI总结 提出首个跨尺度训练与评估范式,通过多倍率视觉问答任务增强病理视觉语言模型的跨尺度推理能力,并构建高质量基准数据集Scale-VQA及模型ScaleReasoner-R1,实现最优性能。
Comments MICCAI 2026
不平衡下的遗忘:多模态大语言模型遗忘中的公平性基准测试
机构 * University of Trento(特伦托大学) ; Fondazione Bruno Kessler(布鲁诺·凯斯勒基金会)
专题命中 视觉问答 :visual question answering(abstract);multimodal large language model(abstract);MLLM(abstract_cn);分类 cs.CV、cs.AI
AI总结 研究多模态大语言模型遗忘中的公平性问题,提出FAIRGET基准和FAUN算法,通过模拟现实场景下不平衡的遗忘请求,在考虑数据不平衡性质时遗忘身份,实验证明该方法在遗忘质量和公平性上具有优越性。
Comments 33 pages
D3VL:利用语言模型从3D时间序列数据和视频中理解驾驶场景
机构 * Bradley Department of Electrical and Computer Engineering, Virginia Tech(弗吉尼亚理工大学布拉德利电气与计算机工程系) ; Virginia Tech Transportation Institute(弗吉尼亚理工大学交通研究所) ; Sanghani Center for Artificial Intelligence and Data Analytics(桑哈尼人工智能与数据分析中心)
专题命中 视觉问答 :MLLM(abstract,abstract_cn);multimodal large language model(abstract);分类 cs.CV、cs.AI
AI总结 本文针对自动驾驶中多模态大语言模型,提出D3VL框架,整合2D和3D时间序列数据,回答交通场景相关问题,在KITTI问答数据集上性能提升11%,并引入Waymo QA数据集扩展以评估模型在多样驾驶条件下处理3D和时间序列数据的能力。
Comments Accepted to IEEE IV 2026