RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Real-Time Multi-Modal and Cross-Scale 3D Reconstruction and Perception for Intelligent Digital Agriculture = 지능형 디지털 농업을 위한 실시간 다중모달·다중스케일 3차원 재구성 및 인지

    한글로보기

    https://www.riss.kr/link?id=T17389346

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Achieving reliable 3D reconstruction and perception in commercial orchards and greenhouse agricultural environments remains a long-standing challenge. This arises from complex structural layouts, dense foliage, seasonal illumination variation, and frequent degradation of GNSS reception under canopy conditions. High-precision localization and multi-scale modeling are essential for digital agriculture applications, including autonomous navigation, fruit phenotyping, yield estimation, and growth monitoring. Traditional single-modality sensing and loosely coupled fusion approaches suffer from drift accumulation, spatial inconsistency, and incomplete representations of orchard and greenhouse geometry. These limitations restrict the operational reliability of robotic platforms and prevent robust decision support in environments with occlusions, narrow passages, and heterogeneous viewpoints. To address these issues, this dissertation developed a unified multi-scale, multi-modal sensing framework for real-time reconstruction, geo-localization, and plant trait estimation across greenhouse rows and orchard-scale farmland. Firstly, a real-time 3D orchard reconstruction approach was developed based on two-sided row alignment and feature-consistent multi-camera odometry. Unified feature extraction and refined pointcloud odometry reduced drift and reinforced geometric consistency, while dual loop-closure detection prevented global trajectory divergence. A pose-graph optimization integrated multiple alignment cues, producing consistent large-scale maps under severe occlusion. Field trials across four apple-orchard rows achieved an APE RMSE below 16 mm and a local distance MAPE below 0.12 %, while enabling real-time operation exceeding 12.5 FPS on embedded hardware. Fruit trait estimation achieved 2.0–2.17 % MAPE for diameter and 5.3–5.6 % for volume. Secondly, greenhouse perception under extreme field-of-view constraints and sparse geometric features was systematically addressed. A tightly coupled RGBD–IMU–LiDAR odometry formulation was introduced, supported by hardware-level synchronization and cross-modal calibration with a thermal camera. Mask2Former expanded sens- ing tolerance to illumination instability, and instance-level fruit segmentation enabled accurate 3D localization for biomass estimation. A modular scheduling scheme ensured real-time processing on embedded computing hardware. Experiments in commercial greenhouse corridors demonstrated an APE RMSE of 1.2 cm without long-term drift. Transformer-based segmentation achieved an F1-score of 0.85 and an IoU of 0.83, while fruit-height estimation achieved R² = 0.804 and a MAPE of 7.0 %. Thirdly, a ground–aerial fusion strategy was presented for orchard-scale 3D reconstruction and geo-localization in GNSS-denied environments by integrating LiDAR-based odometry with UAV. Ground LIO was fused with UAV RSI through a transformer-based dense matching network that learned structural correspondences between ground-level geometric obser- vations and aerial representations. A flow-field decoder estimated cross-view alignment, which was converted into pose constraint factors and integrated into a graph optimization backend. This enabled distortion-free orchard-level mapping tied to aerial geospatial references without GNSS dependence. Multi-season experiments achieved a mean EPE of 3.42 px (≈ 7 cm), with 94.5 % of samples within 5 px, and an APE RMSE of 2.8–3.3 cm. The resulting orchard model enabled accurate tree-height estimation with R² = 0.95 and a MAPE of 3.28 %, supporting geo-spatial digital-twin construction. Overall, these contributions established a scalable perception pipeline spanning local greenhouse rows to orchard-scale outdoor environments. The unified methodology advanced multi-sensor integration, robust optimization, and multi-view structural reasoning for agricultural robotics. The outcomes supported the development of highly accurate digital farmland twins, autonomous navigation systems, and trait-based management tools, enabling data-driven agricultural practices in challenging, real-world field conditions. Keywords: Multi-sensor fusion, Ground-air collaboration, Pose-graph optimization, Trans- former network, Phenotypic estimation, Multi-layer model, Precision agriculture.
    번역하기

    Achieving reliable 3D reconstruction and perception in commercial orchards and greenhouse agricultural environments remains a long-standing challenge. This arises from complex structural layouts, dense foliage, seasonal illumination variation, and fre...

    Achieving reliable 3D reconstruction and perception in commercial orchards and greenhouse agricultural environments remains a long-standing challenge. This arises from complex structural layouts, dense foliage, seasonal illumination variation, and frequent degradation of GNSS reception under canopy conditions. High-precision localization and multi-scale modeling are essential for digital agriculture applications, including autonomous navigation, fruit phenotyping, yield estimation, and growth monitoring. Traditional single-modality sensing and loosely coupled fusion approaches suffer from drift accumulation, spatial inconsistency, and incomplete representations of orchard and greenhouse geometry. These limitations restrict the operational reliability of robotic platforms and prevent robust decision support in environments with occlusions, narrow passages, and heterogeneous viewpoints. To address these issues, this dissertation developed a unified multi-scale, multi-modal sensing framework for real-time reconstruction, geo-localization, and plant trait estimation across greenhouse rows and orchard-scale farmland. Firstly, a real-time 3D orchard reconstruction approach was developed based on two-sided row alignment and feature-consistent multi-camera odometry. Unified feature extraction and refined pointcloud odometry reduced drift and reinforced geometric consistency, while dual loop-closure detection prevented global trajectory divergence. A pose-graph optimization integrated multiple alignment cues, producing consistent large-scale maps under severe occlusion. Field trials across four apple-orchard rows achieved an APE RMSE below 16 mm and a local distance MAPE below 0.12 %, while enabling real-time operation exceeding 12.5 FPS on embedded hardware. Fruit trait estimation achieved 2.0–2.17 % MAPE for diameter and 5.3–5.6 % for volume. Secondly, greenhouse perception under extreme field-of-view constraints and sparse geometric features was systematically addressed. A tightly coupled RGBD–IMU–LiDAR odometry formulation was introduced, supported by hardware-level synchronization and cross-modal calibration with a thermal camera. Mask2Former expanded sens- ing tolerance to illumination instability, and instance-level fruit segmentation enabled accurate 3D localization for biomass estimation. A modular scheduling scheme ensured real-time processing on embedded computing hardware. Experiments in commercial greenhouse corridors demonstrated an APE RMSE of 1.2 cm without long-term drift. Transformer-based segmentation achieved an F1-score of 0.85 and an IoU of 0.83, while fruit-height estimation achieved R² = 0.804 and a MAPE of 7.0 %. Thirdly, a ground–aerial fusion strategy was presented for orchard-scale 3D reconstruction and geo-localization in GNSS-denied environments by integrating LiDAR-based odometry with UAV. Ground LIO was fused with UAV RSI through a transformer-based dense matching network that learned structural correspondences between ground-level geometric obser- vations and aerial representations. A flow-field decoder estimated cross-view alignment, which was converted into pose constraint factors and integrated into a graph optimization backend. This enabled distortion-free orchard-level mapping tied to aerial geospatial references without GNSS dependence. Multi-season experiments achieved a mean EPE of 3.42 px (≈ 7 cm), with 94.5 % of samples within 5 px, and an APE RMSE of 2.8–3.3 cm. The resulting orchard model enabled accurate tree-height estimation with R² = 0.95 and a MAPE of 3.28 %, supporting geo-spatial digital-twin construction. Overall, these contributions established a scalable perception pipeline spanning local greenhouse rows to orchard-scale outdoor environments. The unified methodology advanced multi-sensor integration, robust optimization, and multi-view structural reasoning for agricultural robotics. The outcomes supported the development of highly accurate digital farmland twins, autonomous navigation systems, and trait-based management tools, enabling data-driven agricultural practices in challenging, real-world field conditions. Keywords: Multi-sensor fusion, Ground-air collaboration, Pose-graph optimization, Trans- former network, Phenotypic estimation, Multi-layer model, Precision agriculture.

    더보기

    목차 (Table of Contents)

    • Acknowledgement i
    • Contents ii
    • List of Figures vii
    • List of Tables xv
    • Abbreviations xvi
    • Acknowledgement i
    • Contents ii
    • List of Figures vii
    • List of Tables xv
    • Abbreviations xvi
    • Abstract 1
    • 1. Introduction 3
    • 1.1. Background 3
    • 1.2. Objectives 7
    • 2. Fundamentals 12
    • 2.1. Mathematical and geometric foundations 13
    • 2.1.1. Coordinate frames and homogeneous transformations 13
    • 2.1.2. Lie algebra for 3D motion representation 13
    • 2.2. Multi-modal sensing principles 14
    • 2.2.1. LiDAR range measurement and pointcloud geometry 16
    • 2.2.2. RGBD camera projection and depth uncertainty 17
    • 2.2.3. IMU motion model and pre-integration 19
    • 2.2.4. Complementary properties and sensor synchronization 19
    • 2.3. State estimation for multi-Modal fusion 21
    • 2.3.1. Front-end odometry formulation 21
    • 2.3.2. Tightly-coupled and loosely-coupled fusion architectures 22
    • 2.3.3. Factor graph and MAP estimation 24
    • 2.3.4. Loop-closure and global consistency 25
    • 2.3.5. Optimization strategies based on GTSAM 27
    • 3. Literature Review 29
    • 3.1. Overview and perspective 29
    • 3.2. Structural sources of agricultural environments 30
    • 3.3. Limits of temporal integration and loop-closure mechanisms 30
    • 3.4. Fragility of absolute positioning and GNSS-based assumptions 31
    • 3.5. Cross-scale observability and system-level design implications 32
    • 4. Real-Time 3D Orchard Reconstruction with Robust Two-Sided Row Alignment 34
    • 4.1. Summary 34
    • 4.2. Introduction 35
    • 4.3. Materials and methods 38
    • 4.3.1. Overview of the multi-RGBD SLAM framework 38
    • 4.3.2. Design of the vertically stacked multi-RGBD system 39
    • 4.3.2.1. Extrinsic calibration of the multi-RGBD system 42
    • 4.3.3. Hybrid motion estimation 44
    • 4.3.3.1. Unified multi-view feature extraction 44
    • 4.3.3.2. Multi-view visual odometry estimation 45
    • 4.3.3.3. Pointcloud odometry refinement 47
    • 4.3.3.4. Keyframe decision and apriltag pose constraints 48
    • 4.3.4. Pose-graph construction and optimization 49
    • 4.3.4.1. Dual loop-closure detection 50
    • 4.3.4.2. Incremental and batch pose-graph optimization strategy 52
    • 4.3.5. Module management and runtime scheduling 52
    • 4.3.6. Apple trait extraction 54
    • 4.3.7. Field experimental setup 55
    • 4.3.8. Validation and statistical analysis 57
    • 4.4. Results and discussion 58
    • 4.4.1. Reconstruction overview 58
    • 4.4.2. Global trajectory accuracy 60
    • 4.4.3. Local ground-marker distance accuracy 62
    • 4.4.4. Pose-graph structure and loop-closure statistics 64
    • 4.4.5. Computation performance 66
    • 4.4.6. Ablation study: impact of loop-closure factor density 67
    • 4.4.7. Ablation study: impact of multi-camera configuration 69
    • 4.4.8. Accuracy of reconstructed apple trait estimation 70
    • 4.5. Conclusions 73
    • 5. Multi-Sensor Fusion for Reliable 3D Reconstruction under FOV and Feature
    • Scarcity in Greenhouse Rows 75
    • 5.1. Summary 75
    • 5.2. Introduction 76
    • 5.3. Materials and methods 79
    • 5.3.1. Overall framework 79
    • 5.3.2. Multi-modal sensing platform 80
    • 5.3.2.1. RGB-LiDAR extrinsic calibration 82
    • 5.3.2.2. RGB-thermal extrinsic calibration 84
    • 5.3.3. Tightly-coupled RGBD–IMU–LiDAR odometry 86
    • 5.3.4. Fruit instance segmentation and 3D projection 88
    • 5.3.5. Biomass estimation 90
    • 5.3.6. Performance evaluation metrics 92
    • 5.4. Results and discussion 93
    • 5.4.1. Global trajectory accuracy 93
    • 5.4.2. Pose-graph and loop-closure statistics 95
    • 5.4.3. Time synchronization 96
    • 5.4.4. Fruit segmentation and 3D localization 98
    • 5.4.5. Accuracy of trait estimation 100
    • 5.5. Conclusions 102
    • 6. Ground–Aerial Image Fusion for Orchard 3D Reconstruction in GNSS-Denied
    • Environments 104
    • 6.1. Summary 104
    • 6.2. Introduction 105
    • 6.3. Materials and methods 108
    • 6.3.1. Scenario overview 108
    • 6.3.2. LIO-RSI fusion system 109
    • 6.3.3. Dataset preprocessing and integration 110
    • 6.3.3.1. Generation and tiling of RSI 110
    • 6.3.3.2. BEV projection of LiDAR pointcloud 111
    • 6.3.3.3. Experimental setup 112
    • 6.3.3.4. Dataset collection and characteristics 113
    • 6.3.4. LIO-RSI fusion matching network 115
    • 6.3.4.1. Feature extraction with DINO-V2 backbone 115
    • 6.3.4.2. Cross-attention matching module 116
    • 6.3.4.3. Multi-scale flow refinement 117
    • 6.3.4.4. Loss functions and training strategy 119
    • 6.3.4.5. Performance evaluation metrics 120
    • 6.3.5. Pose-graph construction and optimization 121
    • 6.3.5.1. Geo-localization constraint factor 121
    • 6.3.5.2. Loop-closure detection and constraint 122
    • 6.3.5.3. Pose-graph optimization strategy 124
    • 6.3.5.4. Performance evaluation metrics 124
    • 6.3.6. GIS-driven spatial information application 125
    • 6.4. Results and discussion 126
    • 6.4.1. Performance of the LIO-RSI fusion matching network 126
    • 6.4.2. Analysis of reconstruction accuracy 128
    • 6.4.3. Analysis of geo-localization accuracy 130
    • 6.4.4. Pose-graph visualization and statistics 131
    • 6.4.5. Runtime and resource utilization 133
    • 6.4.6. GIS-driven multi-layer orchard model 134
    • 6.5. Conclusions 137
    • 7. Overall Conclusions and Future Works 138
    • 7.1. Overall conclusions 138
    • 7.2. Future works 140
    • References 144
    • Abstract (Korean) 158
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼