AI 中文总结
本研究针对国际象棋棋盘识别任务,提出结合ViT编码器与DETR风格解码器的ChessQueries方法,在ChessReD基准上大幅提升性能,还发布了新的挑战性数据集。
AI 中文摘要
国际象棋棋盘识别是指将国际象棋棋盘图像映射为每个方格上棋子信息的任务。目前该任务有两个已建立的基准:ChessCog是合成数据集,ChessReD来自单个棋盘设置的智能手机拍摄图片。我们提出ChessQueries,一种结合ViT编码器与DETR风格解码器的新方法,其性能优于现有方法。在ChessReD基准上,我们将现有技术水平从15.3%提升至99.2%,并在分布外数据集上展现出强大能力。我们的方法在两个数据集上达到了任务饱和,平均每个棋盘仅0.01个错误方格(对比现有技术水平分别为3.4和0.15)。我们还分享了一个新的、更具挑战性的公开数据集,该数据集从顶级国际象棋锦标赛的转播内容中解析得到。代码、模型权重以及SLCC数据将被发布。
英文摘要
Chess board recognition is the task of mapping the image of a chess board to the information of which piece is on which square. So far this task has two established benchmarks: ChessCog is synthetic, and ChessReD comes from smartphone pictures of a single chess board setup. We introduce ChessQueries, a new method combining a ViT encoder with a DETR-style decoder, which outperforms existing methods. On the ChessReD benchmark, we improve the state of the art from 15.3% to 99.2%, and demonstrate strong capabilities on out-of-distribution datasets. Our method saturates the task on the two datasets, with an average 0.01 wrong squares per board (vs. SotA: 3.4 / 0.15 respectively). We also share a new, harder public dataset, parsed from broadcasted top-level chess tournaments. Code, model weights and the SLCC data will be released.