arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 7409 信号源:cs.CV, cs.AI, cs.LG

1. 视觉定位与Grounding 7409 篇

2601.00448 2026-01-05 cs.CL cs.AI 57%

Language as Mathematical Structure: Examining Semantic Field Theory Against Language Games

语言作为数学结构:对语义场理论与语言游戏的检验

Dimitris Vartziotis

机构 * TWT Science & Innovation(TWT科学与创新) NIKI - Digital Engineering(NIKI数字工程)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文通过对比语义场理论与语言游戏,探讨语言的数学结构与社会建构的互补性,提出新的理论指导AI架构的方向。

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.00156 2026-01-05 cs.CV 57%

Focal-RegionFace: Generating Fine-Grained Multi-attribute Descriptions for Arbitrarily Selected Face Focal Regions

聚焦区域面部:为任意选定的面部区域生成细粒度多属性描述

Kaiwen Zheng, Junchen Fu, Songpei Xu, Yaoqing He, Joemon M. Jose, Han Hu, Xuri Ge

机构 * University of Glasgow(格拉斯哥大学) Shandong University(山东大学) Institute of Computing Technology, Chinese Academy of Sciences(中国科学院计算技术研究所)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出Focal-RegionFace模型,通过多阶段微调实现对任意面部区域的细粒度多属性描述生成,提升面部状态分析的准确性和可解释性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2310.18961 2026-01-05 cs.CV 57%

AnomalyCLIP: Object-agnostic Prompt Learning for Zero-shot Anomaly Detection

AnomalyCLIP:面向零样本异常检测的对象无关提示学习

Qihang Zhou, Guansong Pang, Yu Tian, Shibo He, Jiming Chen

机构 * Zhejiang University(浙江大学) Singapore Management University(新加坡管理学院) Harvard University(哈佛大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 AnomalyCLIP通过学习对象无关的文本提示,实现跨不同领域的零样本异常检测,提升异常识别的泛化能力。

Comments Accepted by ICLR 2024

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.24684 2026-01-01 cs.CL cs.AI 57%

R-Debater: Retrieval-Augmented Debate Generation through Argumentative Memory

R-Debater:通过论证记忆的检索增强辩论生成

Maoyuan Li, Zhongsheng Wang, Haoyuan Li, Jiamou Liu

机构 * Wuhan College of Communication(武汉通信学院) Wuhan College of Communication University of Auckland(武汉通信学院奥克兰大学) University of Auckland(奥克兰大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 R-Debater通过整合检索和结构化规划,实现了更忠实、一致且连贯的多轮辩论生成。

Comments Accepteed by AAMAS 2026 full paper

详情

展开后加载摘要…

URL PDF HTML 收藏
2510.22693 2026-01-01 cs.CV 57%

VADTree: Explainable Training-Free Video Anomaly Detection via Hierarchical Granularity-Aware Tree

VADTree: 通过分层粒度感知树实现可解释的无训练视频异常检测

Wenlong Li, Yifei Xu, Yuan Rao, Zhenhua Wang, Shuiguang Deng

机构 * School of Software, Xi’an Jiaotong University(西安交通大学软件学院) China Railway Xi’an Group(中国铁路西安集团) College of Computer Science and Technology, Zhejiang University(浙江大学计算机科学与技术学院)

专题命中 视觉定位与Grounding :visual language model(abstract);分类 cs.CV

AI总结 VADTree通过分层粒度感知树结构实现无训练视频异常检测,利用预训练模型知识和多维先验提升异常感知与推理能力。

Comments NeurIPS 2025 poster

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.23557 2025-12-30 cs.CR cs.AI 57%

Toward Trustworthy Agentic AI: A Multimodal Framework for Preventing Prompt Injection Attacks

迈向可信的代理AI:一种多模态框架用于防止提示注入攻击

Toqeer Ali Syed, Mishal Ateeq Almutairi, Mahmoud Abdel Moaty

机构 * Faculty of Computer and Information System(计算机与信息系统系) Islamic University of Madinah(麦地那伊斯兰大学) Arab Open University-Bahrain(巴林阿拉伯开放大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.AI

AI总结 本文提出一种多模态框架,通过溯源感知机制防止代理AI中的提示注入攻击,提升系统安全性和稳定性。

Comments It is accepted in a conference paper, ICCA 2025 in Bahrain on 21 to 23 December

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22336 2025-12-30 cs.AI cs.CL 57%

Agent2World: Learning to Generate Symbolic World Models via Adaptive Multi-Agent Feedback

Agent2World: 通过自适应多智能体反馈学习生成符号世界模型

Mengkang Hu, Bowei Xia, Yuran Wu, Ailing Yu, Yude Zou, Qiguang Chen, Shijian Wang, Jiarui Jin, Kexin Li, Wenxiang Jiao, Yuan Lu, Ping Luo

机构 * The University of Hong Kong(香港大学) Xiaohongshu Inc.(小红书公司) UESTC(电子科技大学) Harbin Institute of Technology(哈尔滨工业大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 Agent2World通过自适应多智能体反馈学习生成符号世界模型,提升推理和微调性能,实现30.95%的相对提升。

Comments 48 pages, 15 tables, 7 figures, Project page: https://agent2world.github.io

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.22199 2025-12-30 cs.AI 57%

Bidirectional RAG: Safe Self-Improving Retrieval-Augmented Generation Through Multi-Stage Validation

双向RAG:通过多阶段验证实现安全的自我改进检索增强生成

Teja Chinthala

机构 * Independent Researcher(独立研究者)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 双向RAG通过多阶段验证实现安全的自我改进检索增强生成,提高了覆盖率并减少了文档数量,展示了RAG系统在严格验证下的可行性。

Comments 10 pages, 2 figures, 2 tables. 36 experiments across 4 datasets with 3 random seeds. Code available upon request

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.02005 2025-12-30 cs.CV 57%

Learning Visual Affordance from Audio

从音频学习视觉可及性

Lidong Lu, Guo Chen, Zhu Wei, Yicheng Liu, Tong Lu

机构 * Nanjing University(南京大学) China Mobile Communications Company Limited Research Institute(中国移动通信有限公司研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 AVAGFormer通过融合音频和视觉信号,实现了从音频学习视觉可及性的任务,提升了交互区域的识别性能。

Comments 15 pages, 10 figures

详情

展开后加载摘要…

URL PDF HTML 收藏
2412.09614 2025-12-30 cs.CV cs.CL 57%

RAVEL: Rare Concept Generation and Editing via Graph-driven Relational Guidance

RAVEL:通过图驱动的关系引导进行罕见概念生成与编辑

Kavana Venkatesh, Yusuf Dalva, Ismini Lourentzou, Pinar Yanardag

机构 * Virginia Tech(弗吉尼亚理工大学) UIUC(伊利诺伊大学香槟分校)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 RAVEL通过图驱动的关系引导方法,提升罕见概念生成与编辑的可控性和可解释性,适用于长尾领域。

Comments Project Page: https://ravel-diffusion.github.io/

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21237 2025-12-25 cs.CV 57%

SegMo: Segment-aligned Text to 3D Human Motion Generation

SegMo:基于段落的文本到3D人体运动生成

Bowen Dang, Lin Wu, Xiaohang Yang, Zheng Yuan, Zhixiang Chen

机构 * University of Sheffield(谢菲尔德大学) University of Glasgow(格拉斯哥大学) Queen Mary University of London(伦敦大学Queen Mary)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 SegMo通过段落对齐实现文本与3D人体运动的精细生成与对齐,提升生成效果并拓展应用范围。

Comments The IEEE/CVF Winter Conference on Applications of Computer Vision 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21150 2025-12-25 cs.CV 57%

ORCA: Object Recognition and Comprehension for Archiving Marine Species

ORCA:面向档案化海洋物种的对象识别与理解

Yuk-Kwan Wong, Haixin Liang, Zeyu Ma, Yiwei Chen, Ziqiang Zheng, Rinaldi Gotama, Pascal Sebastian, Lauren D. Sparks, Sai-Kit Yeung

机构 * Hong Kong University of Science and Technology(香港科技大学) University of Electronic Science and Technology of China(电子科学与技术大学) Indo Ocean Foundation Project(印度洋基金会项目)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 ORCA提出一个多模态基准数据集,用于提升海洋物种识别与理解的研究水平,通过细粒度视觉和文本注释推动领域内方法学进展。

Comments Accepted by The IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.21065 2025-12-25 cs.RO cs.CV 57%

Language-Guided Grasp Detection with Coarse-to-Fine Learning for Robotic Manipulation

基于粗到细学习的语言引导抓取检测用于机器人操作

Zebin Jiang, Tianle Jin, Xiangtong Yao, Alois Knoll, Hu Cao

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本文提出基于粗到细学习的语言引导抓取检测方法,通过跨模态融合和动态卷积头提升抓取精度与鲁棒性,实验证明其在复杂环境中的有效性。

Comments Submitted to IEEE Journal

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20714 2025-12-25 cs.AI cs.CY cs.HC 57%

From Pilots to Practices: A Scoping Review of GenAI-Enabled Personalization in Computer Science Education

从飞行员到实践:关于生成式AI在计算机科学教育中实现个性化的一次范围综述

Iman Reihanian, Yunfei Hou, Qingquan Sun

机构 * School of Computer Science and Engineering, California State University, San Bernardino, CA 92407 , USA(计算机科学与工程学院,加州州立大学,桑 bernardino,CA 92407,美国)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文通过综述32项研究,探讨生成式AI在计算机科学教育中的个性化应用,提出探索优先的采用框架,强调试点和证据驱动的扩展,以提升学习效果并应对相关风险。

Comments Review article. 23 pages, 7 figures, 8 tables. Published in AI (MDPI), 2026

Journal ref AI 2026, 7(1), Article 6

详情

展开后加载摘要…

URL PDF HTML 收藏
2506.11718 2025-12-25 cs.HC cs.AI 57%

Interaction, Process, Infrastructure: A Unified Framework for Human-Agent Collaboration

交互、过程、基础设施:人机协作的统一框架

Yun Wang, Yan Lu

机构 * Microsoft Research Asia(微软亚洲研究院)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出了一种统一框架,通过整合交互、过程和基础设施,解决人机协作中结构表示和适应性问题。

详情

展开后加载摘要…

URL PDF HTML 收藏
2503.11953 2025-12-25 cs.CV 57%

SPOC: Spatially-Progressing Object State Change Segmentation in Video

SPOC: 视频中空间性推进的对象状态变化分割

Priyanka Mandikal, Tushar Nagarajan, Alex Stoken, Zihui Xue, Kristen Grauman

专题命中 视觉定位与Grounding :VLM(abstract);分类 cs.CV

AI总结 SPOC提出空间性推进的对象状态变化分割任务,通过伪标签和动态约束方法,解决视频中对象变化位置和速度的精确定位问题。

Comments Accepted at WACV 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20451 2025-12-24 cs.CV 57%

Beyond Motion Pattern: An Empirical Study of Physical Forces for Human Motion Understanding

超越运动模式:人体运动理解中物理力的实证研究

Anh Dao, Manh Tran, Yufei Zhang, Xiaoming Liu, Zijun Cui

机构 * Michigan State University(密歇根州立大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 本研究通过引入物理推断的力,提升了人体运动理解在步态识别、动作识别和视频描述中的性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20344 2025-12-24 cs.AI 57%

A DeepSeek-Powered AI System for Automated Chest Radiograph Interpretation in Clinical Practice

基于DeepSeek的AI系统用于临床实践中自动胸部X光解读

Yaowei Bai, Ruiheng Zhang, Yu Lei, Xuhua Duan, Jingfeng Yao, Shuguang Ju, Chaoyang Wang, Wei Yao, Yiwan Guo, Guilin Zhang, Chao Wan, Qian Yuan, Lei Chen, Wenjuan Tang, Biqiang Zhu, Xinggang Wang, Tao Sun, Wei Zhou, Dacheng Tao, Yongchao Xu, Chuansheng Zheng, Huangxuan Zhao, Bo Du

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.AI

AI总结 基于DeepSeek的AI系统在临床实践中实现自动胸部X光解读,提升诊断可靠性与效率,减少解读时间并获得专家认可。

Comments arXiv admin note: substantial text overlap with arXiv:2507.19493

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20278 2025-12-24 cs.AI 57%

Synthesizing Procedural Memory: Challenges and Architectures in Automated Workflow Generation

生成程序记忆:自动化工作流生成中的挑战与架构

Nishant Gaurav, Adit Akarsh, Ankit Ranjan, Manoj Bajaj

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出通过科学方法自主生成工作流技能,解决自动化技能生成中的四个结构性瓶颈。

Comments 7 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.20257 2025-12-24 cs.CV 57%

LADLE-MM: Limited Annotation based Detector with Learned Ensembles for Multimodal Misinformation

LADLE-MM:基于有限标注的多模态虚假信息检测器

Daniele Cardullo, Simone Teglia, Irene Amerini

机构 * Sapienza University of Rome(罗马萨皮恩扎大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 LADLE-MM是一种基于有限标注的多模态虚假信息检测器,通过学习的集成方法在有限资源下实现高效检测,优于现有方法。

详情

展开后加载摘要…

URL PDF HTML 收藏
2508.12711 2025-12-24 cs.CV 57%

Drifting Away from Truth: GenAI-Driven News Diversity Challenges LVLM-Based Misinformation Detection

远离真相:生成式AI驱动的新闻多样性对基于LVLM的虚假信息检测的挑战

Fanxiao Li, Jiaying Wu, Tingchao Fu, Yunyun Dong, Bingbing Song, Wei Zhou

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文探讨生成式AI驱动的新闻多样性对基于LVLM的虚假信息检测系统造成的挑战,通过DriftBench基准揭示了现有系统在多级漂移下的鲁棒性问题,并指出需要更稳健的解决方案。

详情

展开后加载摘要…

URL PDF HTML 收藏
2511.00908 2025-12-23 cs.CV cs.GR 57%

GraphGeo: Multi-Agent Debate Framework for Visual Geo-localization with Heterogeneous Graph Neural Networks

GraphGeo: 多智能体辩论框架用于具有异构图神经网络的视觉地定位

Heng Zheng, Yuling Shi, Xiaodong Gu, Haochen You, Zijian Zhang, Lubin Gan, Hao Zhang, Wenjun Huang, Jin Huang

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 GraphGeo通过异构图神经网络的多智能体辩论框架,提升视觉地定位的准确性与鲁棒性。

Comments This submission has been withdrawn by the authors due to a fundamental error in the methodology that affects the validity of the main results

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18853 2025-12-23 cs.CV cs.HC 57%

VizDefender: Unmasking Visualization Tampering through Proactive Localization and Intent Inference

VizDefender:通过主动定位和意图推断揭示可视化篡改

Sicheng Song, Yanjie Zhang, Zixin Chen, Huamin Qu, Changbo Wang, Chenhui Li

机构 * East China Normal University(华东师范大学) Hong Kong University of Science and Technology(香港科学与技术大学) School of Computer Science and Technology(计算机科学与技术学院)

专题命中 视觉定位与Grounding :multimodal large language model(abstract);分类 cs.CV

AI总结 VizDefender通过半脆弱水印和意图分析模块,主动定位和推断可视化篡改,有效检测和分析数据篡改与视觉编码篡改。

Comments IEEE Transactions on Visualization and Computer Graphics (IEEE PacificVis'26 TVCG Track)

详情

展开后加载摘要…

URL PDF HTML 收藏
2501.17642 2025-12-23 cs.CV 57%

Efficient Redundancy Reduction for Open-Vocabulary Semantic Segmentation

高效开放词汇语义分割中的冗余减少

Lin Chen, Qi Yang, Kun Ding, Zhihao Li, Gang Shen, Fei Li, Qiyuan Cao, Shiming Xiang

机构 * State Key Laboratory of Multimodal Artificial Intelligence Systems (MAIS), Institute of Automation, Chinese Academy of Sciences(多模态人工智能系统国家重点实验室(MAIS)、自动化研究所、中国科学院) School of Artificial Intelligence, University of Chinese Academy of Sciences(中国科学院大学人工智能学院) School of Software, Shandong University(山东大学软件学院) China Tower Corporation Limited(中国铁塔股份有限公司)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 本文提出ERR-Seg,通过减少冗余信息和优化序列建模,提升开放词汇语义分割的效率与性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.18440 2025-12-23 cs.CL cs.AI 57%

An Agentic AI Framework for Training General Practitioner Student Skills

一种用于培训全科医学生技能的代理AI框架

Victor De Marez, Jens Van Nooten, Luna De Bruyne, Walter Daelemans

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 本文提出了一种代理AI框架,用于培训全科医学生技能,通过统一病例生成、角色驱动对话和基于标准的评估,提升医学教育中虚拟模拟患者的效果。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17733 2025-12-22 cs.IR cs.AI 57%

Diversity Recommendation via Causal Deconfounding of Co-purchase Relations and Counterfactual Exposure

通过共购关系和反事实暴露的因果去偏实现多样性推荐

Jingmao Zhang, Zhiting Zhao, Yunqi Lin, Jianghong Ma, Tianjun Wei, Haijun Zhang, Xiaofeng Zhang

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.AI

AI总结 Cadence通过因果去偏共购关系和反事实暴露提升推荐多样性,同时保持准确性。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17730 2025-12-22 cs.CV 57%

AdaptPrompt: Parameter-Efficient Adaptation of VLMs for Generalizable Deepfake Detection

AdaptPrompt: 为通用深度伪造检测高效调整视觉语言模型参数

Yichen Jiang, Mohammed Talha Alam, Sohail Ahmed Khan, Duc-Tien Dang-Nguyen, Fakhri Karray

机构 * University of Waterloo(滑铁卢大学) University of Bergen(卑尔根大学)

专题命中 视觉定位与Grounding :vision-language model(abstract);分类 cs.CV

AI总结 AdaptPrompt通过参数高效迁移学习框架,利用CLIP模型提升深度伪造检测的泛化能力,实现跨领域和少量样本下的高准确率识别。

Comments Under Review

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17661 2025-12-22 cs.RO cs.LG 57%

Vidarc: Embodied Video Diffusion Model for Closed-loop Control

Vidarc:用于闭环控制的具身视频扩散模型

Yao Feng, Chendong Xiang, Xinyi Mao, Hengkai Tan, Zuyue Zhang, Shuhe Huang, Kaiwen Zheng, Haitian Liu, Hang Su, Jun Zhu

机构 * Dept. of Comp. Sci. and Tech., Institute for AI, BNRist Center, THBI Lab, Tsinghua-Bosch Joint ML Center, Tsinghua University(计算机科学与技术系,人工智能研究院,BNRist中心,THBI实验室,清华-博世联合机器学习中心,清华大学)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 Vidarc 提出了一种基于视频扩散的具身模型,通过掩码反向动力学模型实现快速准确的闭环控制,提升了现实部署中的成功率和效率。

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.17331 2025-12-22 cs.CV 57%

SynergyWarpNet: Attention-Guided Cooperative Warping for Neural Portrait Animation

SynergyWarpNet:基于注意力的协同形变用于神经肖像动画

Shihang Li, Zhiqiang Gong, Minming Ye, Yue Gao, Wen Yao

机构 * MoE Key Lab of Artificial Intelligence, AI Institute, Shanghai Jiao Tong University(人工智能联合实验室,人工智能研究院,上海交通大学) Defense Innovation Institute, Academy of Military Science(国防创新研究院,军事科学院) Intelligent Game and Decision Laboratory(智能游戏与决策实验室)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.CV

AI总结 SynergyWarpNet通过基于注意力的协同形变框架,实现了高保真的说话头合成,解决了传统方法在运动转移和几何基础方面的不足。

Comments Submitted to ICASSP 2026

详情

展开后加载摘要…

URL PDF HTML 收藏
2512.16928 2025-12-22 cs.LG cs.DC 57%

Dion2: A Simple Method to Shrink Matrix in Muon

Dion2: 一种简化矩阵压缩的穆恩方法

Kwangjun Ahn, Noah Amsel, John Langford

机构 * Microsoft Research, AI Frontiers(微软研究院,人工智能前沿)

专题命中 视觉定位与Grounding :grounding(abstract);分类 cs.LG

AI总结 Dion2通过稀疏化矩阵更新来简化穆恩优化器的正交化步骤,从而降低计算和通信开销,提升可扩展性。

Comments https://github.com/microsoft/dion/

详情

展开后加载摘要…

URL PDF HTML 收藏