arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~
arXiv 2609.05257cs.AI

计算机视觉中的常识推理:基础、最新进展与未来方向

Commonsense Reasoning in Computer Vision: Foundations, Recent Advancements, and Future Directions

Bahar Uddin Mahmud, Sumit Barua, Guan Yue Hong, Ajay Gupta, Hexu Liu

首次发表
浏览论文内容

中文总结 AI 辅助

本文综述了将常识知识融入计算机视觉任务的最新进展,介绍了相关方法、当前局限及未来研究方向,旨在开发更智能的视觉系统。

中文摘要 AI 辅助

计算机视觉中的常识推理涵盖了视觉数据与上下文知识的融合,这对提升AI对日常场景的理解至关重要。这种理解不仅能改进机器学习模型,还能增强其与人类及环境进行有意义交互的能力。与旨在识别特定图像内对象的基于CNN的传统视觉模型不同,融入常识知识能让模型以更整体的方式解释场景,从而提升其推理对象间关系与动作的空间能力。这种融合不仅能增强对象识别,还能促进对上下文因素的更深入理解,最终在实际应用中实现更精准的预测与交互。本文对将常识知识融入计算机视觉任务的最新进展进行了全面综述,系统回顾了基于知识图谱、场景图、神经符号模型及常识增强型Transformer的方法,概述了当前与数据集偏差、知识不完整性及融合挑战相关的局限,最后强调了跨模态推理、可扩展常识知识注入及神经符号混合架构等前瞻性研究方向,以开发真正智能的视觉系统。

英文摘要

Commonsense reasoning in computer vision encompasses integrating visual data and contextual knowledge, crucial for enhancing AI's understanding of everyday scenarios. This understanding not only improves machine learning models but also enhances their ability to interact meaningfully with humans and the environment. Unlike CNN-based conventional vision models, which are designed to identify objects within a specific image, incorporating commonsense knowledge enables models to interpret scenes in a more holistic manner, thereby improving their spatial ability to reason about relationships among objects and actions. This integration not only enhances object recognition but also facilitates a deeper understanding of the contextual factors, ultimately leading to more precise predictions and interactions in real-world applications. This paper presents a comprehensive survey of recent developments that integrate commonsense knowledge into computer vision tasks. We systematically review approaches based on knowledge graphs, scene graphs, neuro-symbolic models, and commonsense-augmented transformers. We also outline current limitations related to dataset bias, knowledge incompleteness, and integration challenges. Finally, we highlight prospective research trajectories in cross-modal reasoning, scalable commonsense knowledge injection, and neuro-symbolic hybrid architectures to develop truly intelligent visual systems.

发表机构

  • Lander University(兰德大学)
  • Western Michigan University(西密歇根大学)

机构由 AI 辅助整理,请以论文原文为准。

补充信息

↑