Achieving reliable 3D reconstruction and perception in commercial orchards and greenhouse agricultural environments remains a long-standing challenge. This arises from complex structural layouts, dense foliage, seasonal illumination variation, and fre...
Achieving reliable 3D reconstruction and perception in commercial orchards and greenhouse agricultural environments remains a long-standing challenge. This arises from complex structural layouts, dense foliage, seasonal illumination variation, and frequent degradation of GNSS reception under canopy conditions. High-precision localization and multi-scale modeling are essential for digital agriculture applications, including autonomous navigation, fruit phenotyping, yield estimation, and growth monitoring. Traditional single-modality sensing and loosely coupled fusion approaches suffer from drift accumulation, spatial inconsistency, and incomplete representations of orchard and greenhouse geometry. These limitations restrict the operational reliability of robotic platforms and prevent robust decision support in environments with occlusions, narrow passages, and heterogeneous viewpoints. To address these issues, this dissertation developed a unified multi-scale, multi-modal sensing framework for real-time reconstruction, geo-localization, and plant trait estimation across greenhouse rows and orchard-scale farmland. Firstly, a real-time 3D orchard reconstruction approach was developed based on two-sided row alignment and feature-consistent multi-camera odometry. Unified feature extraction and refined pointcloud odometry reduced drift and reinforced geometric consistency, while dual loop-closure detection prevented global trajectory divergence. A pose-graph optimization integrated multiple alignment cues, producing consistent large-scale maps under severe occlusion. Field trials across four apple-orchard rows achieved an APE RMSE below 16 mm and a local distance MAPE below 0.12 %, while enabling real-time operation exceeding 12.5 FPS on embedded hardware. Fruit trait estimation achieved 2.0–2.17 % MAPE for diameter and 5.3–5.6 % for volume. Secondly, greenhouse perception under extreme field-of-view constraints and sparse geometric features was systematically addressed. A tightly coupled RGBD–IMU–LiDAR odometry formulation was introduced, supported by hardware-level synchronization and cross-modal calibration with a thermal camera. Mask2Former expanded sens- ing tolerance to illumination instability, and instance-level fruit segmentation enabled accurate 3D localization for biomass estimation. A modular scheduling scheme ensured real-time processing on embedded computing hardware. Experiments in commercial greenhouse corridors demonstrated an APE RMSE of 1.2 cm without long-term drift. Transformer-based segmentation achieved an F1-score of 0.85 and an IoU of 0.83, while fruit-height estimation achieved R² = 0.804 and a MAPE of 7.0 %. Thirdly, a ground–aerial fusion strategy was presented for orchard-scale 3D reconstruction and geo-localization in GNSS-denied environments by integrating LiDAR-based odometry with UAV. Ground LIO was fused with UAV RSI through a transformer-based dense matching network that learned structural correspondences between ground-level geometric obser- vations and aerial representations. A flow-field decoder estimated cross-view alignment, which was converted into pose constraint factors and integrated into a graph optimization backend. This enabled distortion-free orchard-level mapping tied to aerial geospatial references without GNSS dependence. Multi-season experiments achieved a mean EPE of 3.42 px (≈ 7 cm), with 94.5 % of samples within 5 px, and an APE RMSE of 2.8–3.3 cm. The resulting orchard model enabled accurate tree-height estimation with R² = 0.95 and a MAPE of 3.28 %, supporting geo-spatial digital-twin construction. Overall, these contributions established a scalable perception pipeline spanning local greenhouse rows to orchard-scale outdoor environments. The unified methodology advanced multi-sensor integration, robust optimization, and multi-view structural reasoning for agricultural robotics. The outcomes supported the development of highly accurate digital farmland twins, autonomous navigation systems, and trait-based management tools, enabling data-driven agricultural practices in challenging, real-world field conditions. Keywords: Multi-sensor fusion, Ground-air collaboration, Pose-graph optimization, Trans- former network, Phenotypic estimation, Multi-layer model, Precision agriculture.