arXivDaily arXiv每日学术速递 周一至周五更新

AI 大模型

视觉大模型 / VLM

视觉语言模型、视觉推理、视觉问答、图文理解和视觉 grounding。

共收录 26349 信号源:cs.CV, cs.AI, cs.LG

1. 视觉推理 4513 篇

1812.01880 2018-12-06 cs.CV 57%

Learning to Compose Dynamic Tree Structures for Visual Contexts

Kaihua Tang, Hanwang Zhang, Baoyuan Wu, Wenhan Luo, Wei Liu

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1811.10561 2018-11-27 cs.CL cs.LG cs.SD eess.AS stat.ML 57%

CLEAR: A Dataset for Compositional Language and Elementary Acoustic Reasoning

Jerome Abdelnour, Giampiero Salvi, Jean Rouat

专题命中 视觉推理 :visual question answering(abstract);分类 cs.LG

Comments NeurIPS 2018 Visually Grounded Interaction and Language (ViGIL) Workshop

详情

展开后加载摘要…

URL PDF HTML 收藏
1710.09490 2018-11-15 cs.CV 57%

Complete 3D Scene Parsing from an RGBD Image

Chuhang Zou, Ruiqi Guo, Zhizhong Li, Derek Hoiem

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

Comments Accepted to International Journal of Computer Vision (IJCV), 2018 arXiv admin note: text overlap with arXiv:1504.02437

详情

展开后加载摘要…

URL PDF HTML 收藏
1806.06765 2018-06-19 cs.LG cs.NE q-bio.NC stat.ML 57%

Modularity Matters: Learning Invariant Relational Reasoning Tasks

Jason Jo, Vikas Verma, Yoshua Bengio

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.LG

Comments Modified abstract to fit arXiv character limit

详情

展开后加载摘要…

URL PDF HTML 收藏
1712.07576 2018-06-18 cs.CV 57%

Learning to Act Properly: Predicting and Explaining Affordances from Images

Ching-Yao Chuang, Jiaman Li, Antonio Torralba, Sanja Fidler

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1711.06526 2018-05-29 cs.CV 57%

Multi-Label Zero-Shot Learning with Structured Knowledge Graphs

Chung-Wei Lee, Wei Fang, Chih-Kuan Yeh, Yu-Chiang Frank Wang

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

Comments CVPR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.03067 2018-04-25 cs.AI 57%

Compositional Attention Networks for Machine Reasoning

Drew A. Hudson, Christopher D. Manning

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.AI

Comments Published as a conference paper at ICLR 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1803.11361 2018-04-02 cs.CV 57%

DDRprog: A CLEVR Differentiable Dynamic Reasoning Programmer

Joseph Suarez, Justin Johnson, Fei-Fei Li

专题命中 视觉推理 :visual question answering(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1802.02598 2018-02-09 cs.CV 57%

Generating Triples with Adversarial Networks for Scene Graph Construction

Matthew Klawonn, Eric Heim

专题命中 视觉推理 :visual question answering(abstract);分类 cs.CV

Comments Accepted to AAAI 2018

详情

展开后加载摘要…

URL PDF HTML 收藏
1801.05302 2018-01-17 cs.CV 57%

Benchmark Visual Question Answer Models by using Focus Map

Wenda Qiu, Yueyang Xianzang, Zhekai Zhang

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

Comments A group project paper for course CS348. arXiv admin note: text overlap with arXiv:1705.03633 by other authors

详情

展开后加载摘要…

URL PDF HTML 收藏
1707.01932 2017-11-10 cs.RO cs.LG stat.ML 57%

End-to-End Learning of Semantic Grasping

Eric Jang, Sudheendra Vijayanarasimhan, Peter Pastor, Julian Ibarz, Sergey Levine

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.LG

Comments 14 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
1504.02437 2017-08-21 cs.CV 57%

Predicting Complete 3D Models of Indoor Scenes

Ruiqi Guo, Chuhang Zou, Derek Hoiem

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1706.01427 2017-06-06 cs.CL cs.LG 57%

A simple neural network module for relational reasoning

Adam Santoro, David Raposo, David G. T. Barrett, Mateusz Malinowski, Razvan Pascanu, Peter Battaglia, Timothy Lillicrap

专题命中 视觉推理 :visual question answering(abstract);分类 cs.LG

详情

展开后加载摘要…

URL PDF HTML 收藏
1612.08153 2016-12-28 cs.CV cs.CG 57%

EgoReID: Cross-view Self-Identification and Human Re-identification in Egocentric and Surveillance Videos

Shervin Ardeshir, Sandesh Sharma, Ali Broji

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1608.08072 2016-08-30 cs.AI 57%

A Novel Approach to Multimedia Ontology Engineering for Automated Reasoning over Audiovisual LOD Datasets

Leslie F. Sikos

专题命中 视觉推理 :grounding(abstract);分类 cs.AI

详情

展开后加载摘要…

URL PDF HTML 收藏
1605.05462 2016-05-19 cs.CV 57%

Dual Local-Global Contextual Pathways for Recognition in Aerial Imagery

Alina Marcu, Marius Leordeanu

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1506.04366 2015-06-16 cs.AI 57%

Artificial general intelligence through recursive data compression and grounded reasoning: a position paper

Arthur Franz

专题命中 视觉推理 :grounding(abstract);分类 cs.AI

Comments 27 pages, 3 figures, position paper

详情

展开后加载摘要…

URL PDF HTML 收藏
1503.06813 2015-04-14 cs.CV 57%

Factorization of View-Object Manifolds for Joint Object Recognition and Pose Estimation

Haopeng Zhang, Tarek El-Gaaly, Ahmed Elgammal, Zhiguo Jiang

专题命中 视觉推理 :visual reasoning(abstract);分类 cs.CV

详情

展开后加载摘要…

URL PDF HTML 收藏
1404.3301 2014-04-15 cs.AI 57%

Efficient Inference and Learning in a Large Knowledge Base: Reasoning with Extracted Information using a Locally Groundable First-Order Probabilistic Logic

William Yang Wang, Kathryn Mazaitis, Ni Lao, Tom Mitchell, William W. Cohen

专题命中 视觉推理 :grounding(abstract);分类 cs.AI

Comments arXiv admin note: substantial text overlap with arXiv:1305.2254

详情

展开后加载摘要…

URL PDF HTML 收藏
2505.05464 2025-07-16 cs.CL 56%

Bring Reason to Vision: Understanding Perception and Reasoning through Model Merging

Shiqi Chen, Jinghan Zhang, Tongyao Zhu, Wei Liu, Siyang Gao, Miao Xiong, Manling Li, Junxian He

机构 * City University of Hong Kong(香港城市大学) Hong Kong University of Science(香港科学大学) National University of Singapore(新加坡国立大学) Northwestern University(西北大学)

专题命中 视觉推理 :vision-language model(abstract);VLM(comments)

Comments ICML 2025. Camera-ready version updated. Our code is publicly available at https://github.com/shiqichen17/VLM_Merging

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.23154 2026-08-25 cs.IR 新提交 50%

The Disconnect Between Better Descriptive Reasoning Trace Quality and Recommendation Effectiveness

更优描述性推理轨迹质量与推荐效果之间的脱节

Gustavo Penha, Juan Elenter, Claudia Hauff, Hugues Bouchard, Paul Bennett, Mounia Lalmas

专题命中 视觉推理 :grounding(abstract)

AI总结 本研究通过2×2因子实验对比语义ID与自然语言标题的推理轨迹质量,发现提升描述性推理轨迹质量无法持续改善传统离线推荐效果。

Comments Accepted at the Recsys'26 Workshop on Agentic and Generative AI for E-Commerce

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22990 2026-08-25 cs.RO 新提交 50%

InstructMove: A Text-Indispensable Benchmark for Instruction-Following Manipulation

InstructMove:一种指令跟随操作的文本不可或缺基准

Mengao Zhao, Ziang Li, Chaodong Huang, Mengchen Ma, Haoyi Jiang, Yiwei Jin, Xinjie Wang, Yun Du, Xuewu Lin, Taojun Ding, Hongyu Xie, Jackson Jiang, Chunlei Yu, Kaihua Zhang, Lichao Huang, Liu Liu, Tianwei Lin, Zhizhong Su

机构 * Horizon Robotics(地平线机器人) WuwenAI(悟文智能) Southeast University(东南大学) Huazhong University of Science and Technology(华中科技大学)

专题命中 视觉推理 :grounding(abstract)

AI总结 针对现有操作基准无法充分检验机器人指令跟随能力的问题,提出文本不可或缺的InstructMove基准,将指令跟随分解为多环节,实验表明其可诊断视觉捷径且模拟数据能提升现实操作性能。

Comments 22 pages

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.22035 2026-08-25 cs.RO 新提交 50%

Ludi${}_{\scriptscriptstyle 0.1}$: An Agentic System for Socially Intelligent Robots

Ludi₀.₁:面向社交智能机器人的智能体系统

Wooseong Chung, William Cong, Jakub Dworakowski, Ethan Ewer, Tri Wahyu Guntara, Yeonwoo Jeong, Tianchong Jiang, Chaewon Kim, Hyunseo Kim, Jinwoo Kim, Jinyeon Kim, Yea-Seul Kim, Jack Kunde, Kangwook Lee, Sangheon Lee, Robert Nowak, Junha Roh

机构 * Ludo Robotics(乐动机器人)

专题命中 视觉推理 :vision-language model(abstract)

AI总结 该研究提出Ludi₀.₁智能体系统,集成多技能与微调视觉语言模型,为社交智能机器人提供可行的人机协作方案,还可生成相关交互轨迹用于开发更融合的机器人基础模型。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.11758 2026-08-13 cs.CL 新提交 50%

AWARe: Mitigating Catastrophic Forgetting via Activation-Weighted Adaptive REtention

AWARe:基于激活加权的自适应保留缓解灾难性遗忘

Juncheng Liao, Jinfan Lv, Guoming Wang, Jupeng Zheng, Ling Xiao, Siliang Tang

机构 * School of Software Technology, Zhejiang University(浙江大学软件学院) College of Intelligent Robotics and Advanced Manufacturing, Fudan University(复旦大学智能机器人与先进制造学院) School of Artificial Intelligence, Sun Yat-Sen University(中山大学人工智能学院) Graduate School of Information Science, Hokkaido University(北海道大学信息科学研究院)

专题命中 视觉推理 :multimodal large language model(abstract)

AI总结 针对多模态大语言模型微调时的灾难性遗忘问题,提出无需修改架构的AWARe方法,通过激活加权控制参数更新,在保留上游能力的同时提升下游性能。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.05825 2026-08-10 cs.CL 版本更新 50%

MoCA: Implicit Social Context Analysis

MoCA:隐式社会语境分析

Wenhao Xu, Kaiwen Zhang, Hao Li, Maowei You, Yongzheng Ji, Siyuan Zuo, Jingxuan Yu, Sina A, Xinyao Tan, Bobo Li, Hao Fei, Mong-Li Lee, Wynne Hsu

机构 * National University of Singapore(新加坡国立大学) University of Oxford(牛津大学) Wuhan University(武汉大学)

专题命中 视觉推理 :multimodal large language model(abstract)

AI总结 本文提出隐式社会语境分析(MoCA)新任务,构建含3108个实例的基准数据集,提出CoDAR框架提升模型性能,但模型与人类推理仍存差距,凸显隐式社会理解难度。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.06013 2026-08-07 cs.HC 新提交 50%

OneEmo: A Unified Multimodal Reasoning Model for Emotion Perception, Understanding, and Interaction

OneEmo:用于情感感知、理解与交互的统一多模态推理模型

Jiahao Huang, Zheng Lian, Jingyi Zhang, Zhide Chen, Xiaojiang Peng, Shaonan Wang

专题命中 视觉推理 :multimodal large language model(abstract)

AI总结 针对现有情感智能多模态模型忽略任务协同的问题,研究人员提出统一多模态情感模型OneEmo,构建EmoWorld-130K数据集并采用Emo-Chord强化学习策略,其性能优于同规模基线模型,参数远少于商业模型且结果具竞争力。

详情

展开后加载摘要…

URL PDF HTML 收藏
2608.01292 2026-08-04 cs.CL 新提交 50%

CrossLex: A Source-Grounded Benchmark for Cross-Jurisdictional Legal Reasoning in Large Language Models

CrossLex:面向大语言模型跨司法辖区法律推理的、基于法律来源的基准测试集

Xiaocui Yang, Xican Tan, Shoujie Chen, Shihan Xiao, Keke Tong, Xinyu Zhou

专题命中 视觉推理 :grounding(abstract)

AI总结 本文提出基于中、加州、德三国法律来源的CrossLex基准测试集,定义三项任务并提出联合评估指标,发现大语言模型在跨司法辖区法律推理上存在不足,旨在推动相关研究。

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.28987 2026-08-03 cs.SE 新提交 50%

A Formalism-Aware Reward Loop for Handwritten UML-to-PlantUML Generation

面向手写UML到PlantUML生成的形式化感知奖励循环

Mersedeh Sadeghi, Simon Scholz, Adrian Psoch-Bajraktari

专题命中 视觉推理 :vision-language model(abstract)

AI总结 本研究提出形式化感知奖励循环,通过微调视觉-语言模型结合Group Relative Policy Optimisation生成PlantUML,提升了转换质量,为建模评估提供了新方向。

Comments Accepted for publication in the Proceedings of the ACM/IEEE 29th International Conference on Model Driven Engineering Languages and Systems (MODELS 2026). This is the accepted author manuscript

详情

展开后加载摘要…

URL PDF HTML 收藏
2601.15297 2026-07-31 cs.CL 版本更新 50%

AfriEconQA: A Benchmark for Quantitative and Temporal Reasoning over World Bank Economic Reports

AfriEconQA:基于世界银行报告的非洲经济分析基准数据集

Edward Ajayi, Mustapha Alaba, David Stephen

机构 * Carnegie Mellon University Africa(卡内基梅隆大学非洲分校)

专题命中 视觉推理 :grounding(abstract)

AI总结 AfriEconQA是一个基于世界银行报告的非洲经济分析基准数据集,旨在测试信息检索和RAG系统在处理复杂经济查询时的性能。

Comments Dataset Explorer: https://afrieconqa.pages.dev/

详情

展开后加载摘要…

URL PDF HTML 收藏
2607.20833 2026-07-30 cs.CL 版本更新 50%

REFACT: Adaptive Fact Restatement for Compact and Faithful Chain-of-Thought Reasoning

REFACT:用于紧凑且忠实的思维链推理的自适应事实重述

Zhensheng Jin, Xin Dai, Zhenghao Liu, Chaojun Xiao, Huiyuan Xie, Yu Gu, Ge Yu, Maosong Sun

机构 * Northeastern University(东北大学) Tsinghua University(清华大学)

专题命中 视觉推理 :grounding(abstract)

AI总结 研究复杂任务中语言模型推理轨迹易偏离上下文的问题,提出REFACT自适应事实重述框架,经两阶段优化,实验证明其能提升长上下文问答等能力,减少令牌消耗,保留更多证据使推理轨迹更优。

详情

展开后加载摘要…

URL PDF HTML 收藏