面向四旋翼无人机的基于稀疏场景流的无IMU机体坐标系状态估计
IMU-Free Body-Frame State Estimation with Sparse Scene Flow for Quadcopters
AI总结:
该研究提出一种无IMU的四旋翼仅视觉状态估计系统,采用扩展卡尔曼滤波与4视图光束平差法,生成机体坐标系状态估计与稀疏场景流,无需GPS等世界坐标系基础设施。
AI中文摘要:
我们提出了一种仅依赖视觉的状态估计系统,适用于配备标准立体相机对且无惯性传感器的X构型四旋翼无人机。该系统完全在机体坐标系中运行,仅需同步的立体图像和电机推力指令。复合流形状态〈SE(3), R³, …〉上的连续-离散扩展卡尔曼滤波器,以静止场景点作为隐式惯性参考,维护机体坐标系下的位姿、速度、角速度、重力及扰动的估计值。特征点采用FAST、Shi-Tomasi算法检测,通过SSD、Lucas-Kanade算法进行时间跟踪,并在相机间通过NCC算法匹配,搜索区域由滤波器推导的位姿和点不确定性预测。基于归一化新息的卡方门限仅允许静止点进入滤波器。该系统还生成携带每点位置、速度和联合协方差的稀疏3D点云,这些点云来自4视图(两个时间戳下的两个立体对)的全光束平差法,该方法以滤波器推导的相对位姿为先验,从立体视差和时间视差联合估计位置和速度。扩展卡尔曼滤波器中的特征点不进入求解器,其信息通过位姿先验体现。点云密度具有空间自适应性:外部焦点引导点分配,在关注区域生成密集覆盖,其余区域则为稀疏覆盖。输出包括机体坐标系状态估计、校准后的位姿变化和稀疏场景流,旨在作为下游以当前机体坐标系为锚定的世界模型的测量源,不依赖GPS、IMU或任何世界坐标系基础设施,不过该架构可兼容未来对这些组件的集成。
英文摘要:
We present a vision-only state estimation system for X-configuration quadcopters equipped with a canonical stereo camera pair and no inertial sensors. The system operates entirely in the body frame, requiring only synchronised stereo images and motor thrust commands. A continuous-discrete extended Kalman filter on a composite manifold state $\langle SE(3), \mathbb{R}^3, \ldots \rangle$ maintains estimates of body-frame pose, velocity, angular velocity, gravity, and disturbances, using stationary scene points as implicit inertial references. Feature points are detected (FAST, Shi-Tomasi), tracked temporally (SSD, Lucas-Kanade) and matched across cameras (NCC), with search regions predicted from filter-derived pose and point uncertainty. Chi-squared gating on the normalised innovation admits only stationary points to the filter. The system also produces a sparse 3D point cloud carrying per-point position, velocity and joint covariance. These come from a 4-view (two stereo pairs at two timestamps) full bundle adjustment that jointly estimates position and velocity from stereo disparity and temporal parallax, with the filter-derived relative pose as a prior. Feature points in the EKF do not enter the solver; their information is reflected through the pose prior. Point cloud density is spatially adaptive: an external focus point directs allocation, producing dense coverage in the region of attention and sparse coverage elsewhere. The output is a body-frame state estimate, a calibrated pose change, and a sparse scene flow. It is intended as a measurement source for a downstream world model anchored in the current body frame, without dependence on GPS, IMU, or any world-frame infrastructure, though the architecture accommodates their future integration.