通过投影卷绕序的帧同步手势检测
Frame-Synchronous Hand Gesture Detection by Projected Winding Order
浏览论文内容
中文总结 AI 辅助
针对手绕长轴旋转的同步检测,提出基于投影卷绕序的过零判据,无需训练数据,在240个近失序列中实现95%检测率、平均误差6.7毫秒。
中文摘要 AI 辅助
视频中的手势识别通常被设定为分类问题:对每一帧进行标注,然后根据标注采取行动。这对于控制场景是足够的,因为命令可以延迟几帧执行而用户不会察觉;但对于同步场景则不足,因为输出必须与手势实际发生的帧对齐。我们针对一种常见动作——张开的手绕其长轴旋转——的同步问题,证明了该问题存在一个精确解,无需分类器、无需训练数据、也无需校准。设s为手腕处两条手掌边缘与外指关节的归一化二维叉积,在相机已执行的投影下。我们证明s可分解为k(θ)cos(θ),且|k|>0处处成立,因此当手掌侧对相机时s恰好为零,其符号跟踪呈现给相机的面。因此检测是单个标量的过零,产生的是瞬时而非区间,并且我们证明该判据对图像镜像、手部尺度和惯用手不变,所有这些都来自其代数形式而非地标估计器。双手交互使用四个指尖作为窗口,观察同一场景的重新样式化版本。我们给出了一个覆盖谓词,在双手交叉时保持正确,而通常的三角剖分则不然,并从噪声统计而非目视检查中推导出三个样式化操作符的每个自由参数。在具有精确真值的数据集上,该判据在240个近失序列中检测到95%的翻转且无假阳性,并将每个翻转定位在发生时刻平均6.7毫秒内,这是手部观察间隔的六分之一,而报告包围样本则为23.0毫秒。比采样更精细地报告事件源于将手势视为连续量的零而非标签。
英文摘要
Gesture recognition on video is normally posed as classification: label each frame, then act on the label. That is adequate for control, where a command may be obeyed several frames late without a user noticing, and inadequate for synchronisation, where an output must be aligned to the frame on which the gesture physically occurred. We take the synchronisation problem for one common movement, the rotation of an open hand about its long axis, and show that it admits an exact solution needing no classifier, no training data and no calibration. Let s be the normalised two-dimensional cross product of the two palm edges at the wrist and the outer knuckles, under the projection the camera already performs. We prove that s factorises as k(theta)cos(theta) with |k| > 0 everywhere, so s vanishes exactly when the palm is edge-on and its sign tracks the face presented to the camera. Detection is therefore a zero crossing of one scalar, which yields an instant rather than an interval, and we prove the criterion invariant to image mirroring, hand scale and handedness, all from its algebraic form rather than from the landmark estimator. A two-handed interaction uses four fingertips as a window onto a restyled version of the same scene. We give a coverage predicate that stays correct when the hands cross, where the usual triangulation does not, and derive each free parameter of the three stylisation operators from a noise statistic rather than by inspection. Against a corpus with exact ground truth the criterion detects 95% of flips with no false positive in 240 near-miss sequences, and places each within 6.7 ms on average of the instant it occurred, a sixth of the interval at which the hand is observed, against 23.0 ms for reporting the bracketing sample. Reporting an event more finely than one samples follows from treating a gesture as the zero of a continuous quantity rather than a label.