arXivDaily arXiv每日学术速递 周一至周五更新
arXiv周末暂无论文更新,休息一下吧,周末愉快~~

基于R-CNN的棋盘局面识别

R-CNN-Based Chess Position Recognition

Paras Govind, Ognjen Arandjelović

arXiv 2610.09191首次发表:更新:

发表机构

University of St Andrews(圣安德鲁斯大学)

机构由 AI 辅助整理,请以论文原文为准。

AI 中文总结

提出基于R-CNN的框架,结合Faster R-CNN棋子识别与Mask R-CNN关键点棋盘检测,通过单应性映射重建国际象棋局面,显著提升检测精度并实现高准确率局面恢复。

AI 中文摘要

仅从三维棋盘的单张图像进行国际象棋局面识别,需要预测棋盘相对于相机的位置和朝向、棋格的占用情况以及棋子类型(包括颜色)。我们提出了一种基于R-CNN的框架,其中包含独立的组件分别用于棋子识别和棋盘几何估计,并将两者的预测结果结合起来以重建局面。对于棋子识别,我们采用Faster R-CNN,并使用类别加权目标和更深的分类头进行改进。该检测器直接作用于输入图像,保留备选的棋子假设,随后利用棋子数量和棋格占用的约束对这些假设进行细化。对于棋盘检测,我们引入了八个带标签的边界关键点构成的八边形排列,使用Mask R-CNN的关键点头进行预测。这些关键点提供了用于单应性估计的冗余对应关系,并编码了棋盘朝向。估计出的单应性将棋子框中的代表点映射到一个8x8网格上。在合成数据集上,对棋子检测的改进使平均精度从61.59%提升至90.14%。预测的棋盘关键点中,有97.11%落在其标注目标相对于图像对角线1%的范围内。使用真实棋子框与预测的单应性相结合,每个测试局面的棋格分配均正确。完整的框架能够精确恢复76.61%的测试局面,并且有96.49%的局面最多只有一个棋格错误。

英文摘要

Performing chess game position recognition solely from a single image of a three-dimensional board requires predicting the position and orientation of the board relative to the camera, the occupancy of squares and the piece type, which includes its colour. We propose an R-CNN-based framework with independent components for piece recognition and board geometry estimation, whose predictions are combined to reconstruct the position. For piece recognition, we adapt Faster R-CNN using a class-weighted objective and a deeper classification head. The detector operates directly on the input image, retaining alternative piece hypotheses that are subsequently refined using constraints on piece counts and square occupancy. For board detection, we introduce an octagonal arrangement of eight labelled boundary keypoints, predicted using the keypoint head of Mask R-CNN. These provide redundant correspondences for homography estimation and encode board orientation. The estimated homography maps representative points from the piece boxes to an 8x8 grid. On a synthetic dataset, the modifications to piece detection increase mean average precision from 61.59% to 90.14%. Of the predicted board keypoints, 97.11% are within 1% of the image diagonal of their labelled targets. Using ground-truth piece boxes with the predicted homographies gives correct square assignments for every test position. The complete framework recovers 76.61% of test positions exactly and 96.49% with at most one incorrect square.

论文原文

arXiv 摘要页 · PDF 原文 · HTML 原文

↑