Accurate alignment between Building Information Modeling (BIM) models and real bridge
scenes is essential for effective augmented reality (AR) applications in inspection, maintenance, and health monitoring. Yet, outdoor bridge environments present se...
Accurate alignment between Building Information Modeling (BIM) models and real bridge
scenes is essential for effective augmented reality (AR) applications in inspection, maintenance, and health monitoring. Yet, outdoor bridge environments present severe challenges for
vision-based tracking, including complex backgrounds, variable illumination, occlusions, and
geometric variability of components. Conventional AR tracking approaches, which rely on sparse feature matching, markers, or coarse geometry, often fail to provide stable and precise registration at the component level, limiting the reliability of BIM overlays for engineering
decision-making.
Deep learning-based image segmentation has the potential to address this limitation by
providing dense, component-level masks that can serve as strong geometric and semantic
constraints for pose estimation. However, existing research typically evaluates segmentation models only on static accuracy metrics and often focuses on damage or generic object classes
rather than bridge components in BIM AR workflows. There is a lack of systematic analysis of
how different segmentation architectures CNN-based semantic segmentation, transformer-
based models, and real-time single-stage detection segmentation networks trade off
segmentation accuracy, inference speed, and mask quality when used as the front end for
BIM-based AR tracking and pose estimation in real bridge inspection scenarios.