This dissertation proposes a novel visual localization framework that enables accurate and reliable local and global pose estimation using only camera-based visual information in a pre-built LiDAR map. Although the PLM provides absolute and dense thre...
This dissertation proposes a novel visual localization framework that enables accurate and reliable local and global pose estimation using only camera-based visual information in a pre-built LiDAR map. Although the PLM provides absolute and dense three-dimensional structural information, aligning camera observations with LiDAR measurements remains inherently challenging due to the fundamental differences between the two sensing modalities. To address this issue, the proposed framework is composed of two complementary modules designed to operate under different levels of drift.
The first module is a plane-based stereo localization system intended for medium-scale drift scenarios on the order of a few meters. This module stabilizes camera pose estimation by leveraging the structural consistency between the visual map and the PLM through the combined use of global planes and surfels. It further increases computational efficiency by triggering the iterative closest point–based registration using drift estimation rather than executing registration at every keyframe. Through PLM-based pose correction, this module effectively eliminates accumulated visual localization drift while maintaining high accuracy and real-time performance.
The second module is a depth-based cross-modal place recognition framework for global initialization, designed for situations in which no prior pose information is available or severe drift makes it impossible to determine the camera's location in the PLM. Camera images and LiDAR scans are transformed into a unified depth representation, allowing the system to utilize Vision Foundation Models with a single shared backbone. Furthermore, a geometry-aware mining strategy constructs reliable training pairs by computing pixel-level, geometry-based overlap scores, substantially improving robustness and performance compared to prior methods.
The proposed framework has been validated extensively across various indoor and outdoor public datasets, as well as simulation environments. Experimental results demonstrate superior accuracy and robustness compared to existing approaches, confirming the effectiveness of the two-module PLM-based visual localization system.