GeoPix: Multi-Modal Large Language Model for Pixel-level Image Understanding in Remote Sensing
专题命中 视觉定位与Grounding :visual question answering(abstract);grounding(abstract);MLLM(abstract);分类 cs.CV
AI 大模型
视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。
专题命中 视觉定位与Grounding :visual question answering(abstract);grounding(abstract);MLLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :visual reasoning(abstract);grounding(abstract);MLLM(abstract);分类 cs.CV
Comments CVPR2025;Code will be released at \url{https://github.com/aim-uofa/SegAgent}
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :vision-language model(abstract);visual question answering(abstract);grounding(abstract);分类 cs.CV
Comments under review. project website: https://github.com/emanuelevivoli/awesome-comics-understanding
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV
Comments Preprint; 41 pages, 32 figures, 16 tables; Project Page at https://drive-bench.github.io/
专题命中 视觉定位与Grounding :VLM(abstract);visual language model(abstract);grounding(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.AI
Comments 11 pages, 8 figures, 3 Tables and 1 Algorithm
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV
Comments Accepted at 2024 European Conference on Computer Vision Workshops (ECCVW). Project page - https://prakashchhipa.github.io/projects/ovod_robustness
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV
Comments 10 pages
Journal ref Robotics: Science and Systems. Semantic Reasoning and Goal Understanding in Robotics 2024
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.AI
Comments Accepted in ICML 2024 MATH AI Workshop
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);multimodal large language model(abstract);分类 cs.CV
Comments CVPR2024 Foundational Few-Shot Object Detection Challenge
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments Accepted by ACL 2024, Main Conference, Long Paper
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments CVPR 2024
专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);grounding(abstract);分类 cs.CV
专题命中 视觉定位与Grounding :grounding(abstract);multimodal large language model(abstract);MLLM(abstract);分类 cs.CV
Comments 30 pages, 10 figures. Code/Project Website: https://github.com/apple/ml-ferret
大语言模型的地理空间概念探测:抽象性、组合性与接地性
机构 * University of Toulouse(图卢兹大学) ; IRIT(信息科学与技术研究院(IRIT))
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
AI总结 该研究针对LLMs的抽象性、组合性与接地性设计测试,构建空间概念基准并开展多模型实验,揭示当前LLMs的概念理解局限,为相关模型的优化提供洞见。
Comments Preprint
MolBioKG:通过多分辨率结构锚定将图外分子锚定到生物医学知识图谱中
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
AI总结 MolBioKG是解决生物医学KG中图外分子冷启动问题的两层系统,通过多分辨率结构锚定实现分子与KG的关联,在多项任务中优于基线并提升关键指标。
Comments Preprint
基于视觉定位的分割基础模型少样本概念提示学习
机构 * Advanced Technology Group(先进技术集团) ; GE HealthCare(GE医疗)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI
AI总结 该研究针对分割基础模型在临床任务中表现不足的问题,提出FS-CPL方法,通过视觉定位学习概念提示,在四个医学影像基准上取得显著性能提升,且与骨干网络无关。
FrED:通过领域知识图谱基础进行外部数据影响估计
机构 * National Centre for Scientific Research “Demokritos”(国家科学研究中心“德谟克利特”) ; University of Glasgow(格拉斯哥大学)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
AI总结 针对生成式AI训练数据归因问题,提出在黑盒设置下运行的概率框架,融合连续特征相似度与领域特定知识图谱,在艺术图像合成和天气预报等领域评估,证明其有效性,为外部数据影响分析提供高效可解释机制。
在布莱克-利特曼模型中锚定投资者观点:神经谓词
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
AI总结 研究在布莱克-利特曼模型下投资组合构建问题,提出用神经谓词作为观点生成机制,将结构化金融分析数据经其处理后映射到模型相关矩阵,方法可解释且完全可微,实现端到端学习。
恢复VLA模型中的语言基础:无需训练的注意力重校准方法
机构 * Tsinghua University(清华大学) ; Singapore Management University(新加坡管理学院) ; Institute of Trustworthy Embodied AI, Fudan University(复旦大学可信具身人工智能研究院) ; Shanghai Key Laboratory of Multimodal Embodied AI(上海多模态具身人工智能重点实验室)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI
AI总结 针对VLA模型在分布外指令下出现“语言盲视”问题,提出无需训练的注意力重校准方法IGAR,通过调整注意力分布恢复语言指令影响,在LIBERO基准和真实机器人上有效减少错误执行。
FirstPass: 在多轮编辑结果中奠定AI科学判断的基础
机构 * The University of Texas at Austin(德克萨斯大学奥斯汀分校) ; RediMinds Inc.(RediMinds公司) ; Disaster Science Operations Center, Western Kentucky University(西肯塔基大学灾害科学运营中心)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
AI总结 针对同行评审AI在领域覆盖、对话建模和评估标准上的不足,提出FirstPass数据集和微调模型,利用Nature Communications多轮评审对话,通过响应损失掩码实现80.5%的编辑结果预测准确率,并生成接近人类长度的评审意见。
Comments Accepted at the AI for Science Workshop at the 43rd International Conference on Machine Learning (ICML 2026). 9 pages, 2 figures, 6 tables
MinhwaNet: 韩国民俗画中忠实但不足的对象定位
机构 * Korea Advanced Institute of Science and Technology (KAIST)(韩国科学技术院)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG
AI总结 提出MinhwaNet,通过部分级检测器生成对象证据图,发现韩国民俗画中符号列表不足以预测画作类型,而符号布局更重要,揭示了忠实但不足的解离现象。
可认证安全RLHF:基于语义基础与固定惩罚约束优化的更安全大语言模型对齐
机构 * Department of Electrical and Computer Engineering(电气与计算机工程系) ; New Jersey Institute of Technology(新泽西理工学院) ; Department of Computer Engineering(计算机工程系) ; Heritage Institute of Technology(遗产理工学院)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG
AI总结 针对现有RLHF方法依赖奖励/成本函数和双变量调优导致性能敏感且缺乏可证明安全保证的问题,提出CS-RLHF,通过语义基础成本模型和固定惩罚约束优化,实现可认证安全对齐,效率提升至少5倍。
面向复杂查询的驾驶视频检索与结构化对齐
机构 * NEC Laboratories, America(美国NEC实验室) ; University of California, Riverside(加州大学河滨分校)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.LG
AI总结 提出STRIVE-D框架,通过弱监督领域视频校准规则、融合视觉语言与关键词检索信号,在驾驶视频检索中实现高达84%的top-1准确率提升。
面向光刻缺陷检测的视觉-语言模型失败感知精炼
机构 * Hanyang University(汉阳大学) ; Korea University(高丽大学) ; Korea Institute of Industrial Technology(韩国生产技术研究院)
专题命中 视觉定位与Grounding :vision-language model(title);分类 cs.CV、cs.AI
AI总结 提出两阶段视觉-语言框架,先微调Qwen3-VL检测缺陷,再通过训练精炼模块修正第一阶段错误,提升检测可靠性。
Comments 6 pages, 3 figures
将评分接地:为可靠视觉语言处理奖励模型的显式视觉前提验证
机构 * Qwen Large Model Application Team, Alibaba(阿里巴巴大模型应用团队) ; Alibaba Institute of Automation, Chinese Academy of Sciences(中国科学院自动化研究所阿里巴巴分所) ; Beijing University of Posts and Telecommunications(北京邮电大学)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI
AI总结 本文提出EVPV方法,通过显式验证视觉前提提升视觉语言处理奖励模型的可靠性,实验表明其在多模态推理基准上提升了重排序准确率。
Comments 27 pages, 4 figures, 10 tables. Evaluated on VisualProcessBench and six multimodal reasoning benchmarks (LogicVista, MMMU, MathVerse-VO, MathVision, MathVista, WeMath). Includes ablations and causal analysis via controlled constraint corruption. Code: https://github.com/Qwen-Applications/EVPV-PRM
基于视觉和语言模型的合成数据生成基础
机构 * Graduate School of Informatics(信息学院)
专题命中 视觉定位与Grounding :grounding(title);分类 cs.CV、cs.AI
AI总结 本文提出一种视觉-语言 grounded 框架,用于可解释的遥感合成数据增强与评估,引入 ARAS400k 数据集,包含 100k 真实图像和 300k 合成图像,用于语义分割和图像描述生成。
Comments Accepted for presentation at IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Synthetic Data for Computer Vision Workshop (SynData4CV) 2026