arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

2026-01-30 至 2026-01-30 共收录 14 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 14 篇

2601.21634 2026-01-30 cs.CV 83%

RSGround-R1: Rethinking Remote Sensing Visual Grounding through Spatial Reasoning

RSGround-R1: 重新思考通过空间推理的遥感视觉定位

Shiqi Huang, Shuting He, Bihan Wen

机构 * School of Electrical and Electronic Engineering, Nanyang Technological University(电气电子工程学院,南洋理工大学) MoE Key Laboratory of Interdisciplinary Research of Computation and Economics, Shanghai University of Finance and Economics(教育部计算与经济交叉学科重点实验室,上海财经大学)

专题命中 视觉定位与Grounding :grounding(title,abstract);multimodal large language model(abstract);分类 cs.CV

AI总结 RSGround-R1通过引入空间推理引导的后训练框架,提升遥感图像中目标物体的定位精度与鲁棒性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21296 2026-01-30 cs.LG cs.AI 76%

Grounding and Enhancing Informativeness and Utility in Dataset Distillation

基于信息性和效用的蒸馏数据集 grounding

Shaobo Wang, Yantai Yang, Guo Chen, Peiru Li, Kaixin Li, Yufa Zhou, Zhaorun Chen, Linfeng Zhang

机构 * EPIC Lab, SJTU(SJTU实验室) Shanghai Jiao Tong University(上海交通大学) National University of Singapore(新加坡国立大学) Duke University(杜克大学) The University of Chicago(芝加哥大学)

专题命中 视觉定位与Grounding :grounding(title);分类 cs.AI、cs.LG

AI总结 本文提出InfoUtil框架,通过博弈论和梯度范数优化,提升数据集蒸馏的信息性和效用,实验显示在ImageNet-1K上性能提升6.1%。

Comments Accepted by ICLR 2026, 20 pages, 9 figures, 11 tables

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22114 2026-01-30 cs.CV cs.AI cs.SY eess.SY 73%

SINA: A Circuit Schematic Image-to-Netlist Generator Using Artificial Intelligence

SINA:一种使用人工智能的电路原理图到网表生成器

Saoud Aldowaish, Yashwanth Karumanchi, Kai-Chen Chiang, Soroosh Noorzad, Morteza Fayazi

机构 * University of Utah, Salt Lake City, UT, USA(犹他大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);VLM(abstract);分类 cs.CV、cs.AI

AI总结 SINA利用深度学习、CCL和OCR技术,结合视觉-语言模型,实现了高准确率的电路原理图到网表自动转换。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21199 2026-01-30 cs.CV cs.AI 73%

Thinker: A vision-language foundation model for embodied intelligence

Thinker:一个用于具身智能的视觉-语言基础模型

Baiyu Pan, Daqin Luo, Junpeng Yang, Jiyuan Wang, Yixuan Zhang, Hailin Shi, Jichao Jiao

专题命中 视觉定位与Grounding :vision-language model(abstract);grounding(abstract);分类 cs.CV、cs.AI

AI总结 Thinker通过构建大规模数据集和改进输入方式,在机器人感知与推理任务中实现了最先进的性能。

Comments IROS 2025, 4 pages, 3 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.22041 2026-01-30 cs.MA cs.AI cs.CV cs.LG 67%

Learning to Communicate Across Modalities: Perceptual Heterogeneity in Multi-Agent Systems

在多智能体系统中学习跨模态交流:感知异质性

Naomi Pitzer, Daniela Mihai

机构 * University of Southampton(索姆塞特大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 研究通过异质多模态交流游戏探讨智能体在感知异质性下的交流机制,发现单模态系统更高效,多模态系统需更多信息交换,位扰动实验揭示了意义的分布编码特性。

Comments To be published in EvoLang XVI proceedings. 15 pages, 17 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2411.13578 2026-01-30 cs.CV cs.AI cs.LG 67%

CMOOD: Concept-based Multi-label OOD Detection

基于概念的多标签异常检测

Zhendong Liu, Yi Nian, Yuehan Qin, Henry Peng Zou, Li Li, Xiyang Hu, Yue Zhao

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV、cs.AI、cs.LG

AI总结 COOD提出了一种基于概念的零样本多标签OOD检测框架,通过增强语义空间和改进评分函数,有效区分复杂标签依赖的OOD样本。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.00590 2026-01-30 cs.CL cs.AI cs.LG 62%

Wikontic: Constructing Wikidata-Aligned, Ontology-Aware Knowledge Graphs with Large Language Models

Wikontic: 构建与维基数据对齐、本体感知的知识图谱

Alla Chepurova, Aydar Bulatov, Mikhail Burtsev, Yuri Kuratov

机构 * Cognitive AI Systems Lab(认知人工智能系统实验室) Moscow Independent Research Institute of Artificial Intelligence(莫斯科独立人工智能研究所) London Institute for Mathematical Sciences(伦敦数学科学研究所)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 Wikontic通过多阶段流程构建维基数据对齐、本体感知的知识图谱,提升KG质量并实现高效构建。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21615 2026-01-30 cs.LG cs.AI 62%

Beyond Parameter Finetuning: Test-Time Representation Refinement for Node Classification

超越参数微调:用于节点分类的测试时间表示精调

Jiaxin Zhang, Yiqi Wang, Siwei Wang, Xihong Yang, Yu Shi, Xinwang Liu, En Zhu

机构 * National University of Defense Technology(国防科技大学) Intelligent Game and Decision Lab(智能游戏与决策实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI、cs.LG

AI总结 本文提出TTReFT框架,通过改进测试时间适应方法,解决图神经网络在分布外测试中的性能下降问题,提升节点分类的准确性和实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2507.13993 2026-01-30 eess.IV cs.AI cs.CV 62%

OrthoInsight: Rib Fracture Diagnosis and Report Generation Based on Multi-Modal Large Models

OrthoInsight:基于多模态大模型的肋骨骨折诊断与报告生成

Ningyong Wu, Jiangbo Zhang, Wenhong Zhao, Jinzhi Wang, Chenzhan Yu, Zhigang Xiu, Duwei Dai, Ziyu Xu, Yongli Yang

机构 * Organizational Management Department, School of Management, Xi’an Jiaotong University(管理学院组织管理部,西安交通大学) West China Longquan Hospital, Sichuan University(四川大学西部临床医学院) School of Electronic Science and Engineering, Xi’an Jiaotong University(西安交通大学电子科学与工程学院) Systems Engineering Institute, Xi’an Jiaotong University(西安交通大学系统工程研究院) Institute of Medical Artificial Intelligence, the Second Affiliated Hospital of Xi’an Jiaotong University(西安交通大学第二附属医院医学人工智能研究所) School of Human Settlements and Civil Engineering, Xi’an Jiaotong University(西安交通大学人居环境与土木工程学院) School of Life Science and Technology, Xi’an Jiaotong University(西安交通大学生命科学与技术学院)

专题命中 视觉定位与Grounding :LLaVA(abstract);分类 cs.CV、cs.AI

AI总结 OrthoInsight通过多模态大模型实现肋骨骨折的自动诊断与报告生成,结合CT图像分析与医学知识图谱,提升诊断效率和临床实用性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.21796 2026-01-30 cs.CL cs.AI 57%

KID: Knowledge-Injected Dual-Head Learning for Knowledge-Grounded Harmful Meme Detection

KID: 基于知识注入的双头学习用于知识引导的有害迷因检测

Yaocong Li, Leihan Zhang, Le Zhang, Qiang Yan

机构 * School of Economics and Management, Beijing University of Posts and Telecommunications(经济管理学院,北京邮电大学) College of Computing, Beijing Information Science and Technology University(计算机学院,北京信息科技大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 KID通过知识注入和双头学习框架,提升有害迷因检测的准确性与泛化能力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.13518 2026-01-30 cs.CV 57%

Text-driven Online Action Detection

基于文本的在线动作检测

Manuel Benavent-Lledo, David Mulero-Pérez, David Ortiz-Perez, Jose Garcia-Rodriguez

机构 * Department of Computer Technology, University of Alicante(阿拉维大学计算机技术系) ValgrAI - Valencian Graduate School and Research Network of Artificial Intelligence(瓦伦西亚人工智能研究生学校和研究网络) Institute of Informatics Research, University of Alicante(阿拉维大学信息研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出TOAD模型,利用CLIP文本嵌入实现高效的零样本和少样本在线动作检测,其在THUMOS14数据集上的mAP达到82.46%

Comments Published in Integrated Computer-Aided Engineering

Journal ref Integrated Computer-Aided Engineering. 2025;32(4):415-423

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20895 2026-01-30 cs.LG 57%

Faster Predictive Coding Networks via Better Initialization

通过更好的初始化提升预测编码网络的速度

Luca Pinchetti, Simon Frieder, Thomas Lukasiewicz, Tommaso Salvatori

机构 * Department of Computer Science, University of Oxford(牛津大学计算机科学系) Institute of Logic and Computation, Vienna University of Technology(维也纳技术大学逻辑与计算研究所) VERSES AI Research Lab(VERSES AI研究实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 本文提出了一种改进的初始化方法,通过优化预测编码网络的初始化以提升训练效率和性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20420 2026-01-30 cs.LG 57%

Concept Component Analysis: A Principled Approach for Concept Extraction in LLMs

概念成分分析:一种用于大语言模型中概念提取的原则性方法

Yuhang Liu, Erdun Gao, Dong Gong, Anton van den Hengel, Javen Qinfeng Shi

机构 * Australian Institute for Machine Learning, The University of Adelaide(澳大利亚机器学习研究所,阿德莱德大学) School of Computer Science and Engineering, The University of New South Wales(计算机科学与工程学院,新南威尔士大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 本文提出概念成分分析(ConCA)方法,通过无监督线性解混从LLM中提取可解释概念,提供理论支持优于现有SAEs方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.20230 2026-01-30 cs.CL cs.HC 50%

Unit-Based Agent for Semi-Cascaded Full-Duplex Dialogue Systems

基于单元的代理用于半级联全双工对话系统

Haoyuan Yu, Yuxuan Chen, Minjie Cai

机构 * Hunan University(湖南大学) Gongdao Technology(公道科技) Jilin University(吉林大学)

专题命中 视觉定位与Grounding :multimodal large language model(abstract)

AI总结 本文提出基于单元的半级联全双工对话系统,利用多模态大语言模型和辅助模块实现高效对话处理,实验显示其在挑战赛中表现优异。

Comments ICASSP 2026 (Grant Challenge). https://github.com/yu-haoyuan/fd-badcat

详情

展开后加载摘要…

URL PDF HTML 收藏