arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

计算机视觉导论

Introduction to Computer Vision

Stan Birchfield

arXiv 2609.39627首次发表:更新:

发表机构

NVIDIA; University of Washington(英伟达; 华盛顿大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

本书以代码优先方式系统介绍计算机视觉,涵盖经典图像处理、三维视觉和深度学习,每种算法均用Python实现并与库函数校验,旨在为读者提供自成一体的算法理解参考。

AI 中文摘要

本书以代码优先的方式介绍计算机视觉,涵盖经典二维图像处理、经典三维视觉和深度学习。全书分为三部分,共44个短章节,每个主题都从第一性原理出发构建:图像算术与形态学;卷积、金字塔和频域滤波;特征检测、光流和立体视觉;射影几何、相机标定和运动恢复结构;以及现代深度学习的完整脉络,从单个神经元到卷积网络、反向传播、经典架构、迁移学习、目标检测、语义分割和实例分割,最后以混合精度和并行训练等工程考量作结。每种技术都直接用Python和NumPy或PyTorch实现,并与相应的OpenCV或PyTorch库函数进行数值校验,使读者不仅看到数学原理,还能看到其在真实和合成数据上的具体行为。这些材料是在AI辅助下从免费在线课程笔记中提炼而成,将大量工作代码浓缩为简洁的数学阐述,同时全程保留经过验证、可复现的结果。本书旨在为希望理解计算机视觉算法及其Python实现的学生和从业者提供一本自成一体的参考书。

英文摘要

This book presents a code-first introduction to computer vision, spanning classical 2D image processing, classical 3D vision, and deep learning. Organized as 44 short chapters across three parts, the book builds each topic from first principles: image arithmetic and morphology; convolution, pyramids, and frequency-domain filtering; feature detection, optical flow, and stereo; projective geometry, camera calibration, and structure from motion; and the full arc of modern deep learning, from a single neuron through convolutional networks, backpropagation, classic architectures, transfer learning, object detection, and semantic and instance segmentation, concluding with engineering considerations like mixed-precision and parallel training. Every technique is implemented directly in Python and NumPy or PyTorch and checked numerically against the corresponding OpenCV or PyTorch library function, so readers see not just the mathematics but its concrete behavior on real and synthetic data. The material was distilled with AI assistance from freely available online course notes, condensing extensive working code into concise mathematical exposition while preserving verified, reproducible results throughout. It is intended as a self-contained reference for students and practitioners who want to understand computer vision algorithms and their Python implementations.

Comments217 pages. For online notes and code, see https://sbirchfield.github.io/cvintro

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑