본 연구는 토공 중심의 건설 현장에서 생산성을 분석하고 개선하기 위한 통합적인 비전 기반 프레임워크를 제안한다. 본 프레임워크는 설명 가능한 객체 검출, 표준화된 작업 분류, 단안(單...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17314733
서울 : 서울대학교 대학원, 2025
학위논문(박사) -- 서울대학교 대학원 , 건설환경공학부 건설관리 , 2025. 8
2025
영어
624
서울
x, 162 ; 26 cm
지도교수: 지석호
I804:11032-000000192510
0
상세조회0
다운로드본 연구는 토공 중심의 건설 현장에서 생산성을 분석하고 개선하기 위한 통합적인 비전 기반 프레임워크를 제안한다. 본 프레임워크는 설명 가능한 객체 검출, 표준화된 작업 분류, 단안(單...
본 연구는 토공 중심의 건설 현장에서 생산성을 분석하고 개선하기 위한 통합적인 비전 기반 프레임워크를 제안한다. 본 프레임워크는 설명 가능한 객체 검출, 표준화된 작업 분류, 단안(單眼) 카메라 기반 3D 자세 추정 기법을 결합하여, 별도의 고가 센서 없이 장비 운용 성과를 정량적으로 분석할 수 있도록 설계되었다. 특히, Grad-CAM 기반 오류 진단 모듈을 통해 객체 검출 오류 유형(예: 가림, 이상 시점 등)을 체계적으로 분류하고 시각적으로 해석할 수 있으며, 이를 통해 모델 성능 향상 및 사용자 신뢰 확보가 가능하다. 또한, 장비나 작업에 구애받지 않는 일반화된 활동 분류 체계를 수립하여 다양한 시공 장면에서 일관된 작업 분류가 가능하도록 하였다. 단안 카메라 영상만으로 3D 자세를 추정하기 위해, 장비 부분을 잘라낸 이미지 DB와 가상 모델 정합 기술을 활용하였다. 이로써 swing, bucket, boom 등 장비 부위별 세부 동작 분석이 가능해졌고, 공간적 비효율의 정량적 진단이 가능해졌다. 세 개의 실제 건설 현장에서 실험을 수행한 결과, 객체 검출 및 작업 분류 정확도는 최대 99%를 기록하였으며, 3D 자세 추정의 평균 RMSE는 1m 이하로 나타났다. 분석 결과, 작업 지연의 주요 원인은 불균형한 장비 디스패칭 및 과도한 토사 정리 작업 등으로 확인되었다. 이러한 분석 결과를 바탕으로, 시뮬레이션 기반 의사결정 지원 시스템을 구축하여 장비 회전각, 작업 배정, 트럭 경로 등을 조정하며 생산성 향상 시나리오를 실험하였다. 예를 들어, 디스패치 전략을 최적화함으로써 전체 유휴 시간을 30% 이상 단축할 수 있었다.
본 연구는 비전 기반 기술만으로 복잡한 건설 현장의 생산성 분석과 개선 전략 수립이 가능함을 입증하였으며, 향후 다양한 작업 환경과 실시간 시스템 통합으로의 확장을 통해 실무적 활용 가능성을 더욱 높일 수 있다.
다국어 초록 (Multilingual Abstract)
This study proposes an integrated, vision-based framework for productivity analysis and optimization in earthwork construction environments. The approach combines explainable object detection, standardized activity classification, and monocular 3D pos...
This study proposes an integrated, vision-based framework for productivity analysis and optimization in earthwork construction environments. The approach combines explainable object detection, standardized activity classification, and monocular 3D pose estimation to quantify operational performance without relying on costly or invasive sensors. A Grad-CAM–based diagnosis module was developed to identify and classify common detection errors (e.g., occlusion, abnormal viewpoints), enabling targeted retraining and model interpretability. Additionally, a generalized activity system was established to overcome task- and equipment-specific limitations and support consistent activity labeling across diverse operational scenes. Monocular 3D pose estimation was achieved using cropped image databases and virtual model alignment, eliminating the need for depth sensors. The resulting pose data enabled motion-level analysis of equipment such as swing angle, bucket articulation, and arm behavior, offering deeper insight into spatial inefficiencies. The proposed pipeline was validated through experiments on three earthwork sites, demonstrating high object detection and activity classification accuracy (up to 99%) and robust 3D pose estimation with an average RMSE below 1 meter. Productivity imbalances were identified across sites using both 2D and 3D data, revealing factors such as dispatch misalignment and excessive soil management time. Building on these insights, a simulation-based decision support module was developed to model alternative operational strategies. By adjusting controllable elements—such as excavator swing angle, dispatch scheduling, and truck routing—the system quantitatively evaluated potential productivity gains. In one case, dispatch optimization reduced total idle time by over 30%.
This research demonstrates that vision-based methods can not only assess productivity in complex construction environments but also support data-driven operational planning. The proposed framework serves as a scalable, sensor-free alternative for real-time monitoring, diagnosis, and optimization in dynamic jobsite conditions. Future work will extend its applicability to broader construction tasks and explore real-time system integration.
목차 (Table of Contents)