RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Context-Aware and Multimodal Recognition of Operational States for Construction Site Monitoring = 건설 현장 모니터링을 위한 상황인지 및 멀티모달 기반 운영 상태 인식

    한글로보기

    https://www.riss.kr/link?id=T17314537

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    건설 프로젝트의 규모와 복잡성이 커짐에 따라 안전, 효율성, 상황 인식을 향상시킬 수 있는 자동화된 모니터링 시스템의 필요성이 커지고 있다. 최근 컴퓨터 비전 기반 모니터링 기술의 발전으로 작업자, 장비, 자재 등의 탐지 및 활동 인식을 가능하게 했지만, 기존 모델들은 시각적으로 모호하거나 문맥적으로 복잡한 조건에서 자재 및 장비의 작동 상태를 추론하는 데 한계가 있다. 본 논문에서는 비전 기반 모니터링 시스템의 두 가지 근본적인 한계를 다루고자 한다. 첫째, 기존 객체 탐지 모델은 형상, 색상, 질감과 같은 시각적 외형 단서에 크게 의존한다. 이러한 특징은 일반적인 객체 탐지에는 적합하지만, 특히 개별 상태가 유사한 외형을 보이는 경우 자재 및 장비의 작동 상태를 식별하는 데는 부적합하다. 이러한 한계는 건설 환경에서 문맥 의존적인 상황을 자주 오인하게 만들 수 있다. 둘째, 현재 접근 방식은 단일 모달리티 모델에 국한되어 있으며, 깊이 데이터, 소리, 시간적 패턴과 같은 보완적인 문맥 단서 없이 시각 정보만을 활용하는 경우가 많다. 이러한 비전 기반 모니터링 시스템은 가림 현상, 객체 간의 복잡한 상호작용이 많이 발생하는 동적인 건설 환경에서 문맥적 추론을 수행하는데 매우 제한적이다. 이러한 한계를 해결하기 위해 본 논문은 공간 추론 및 멀티 모달리티 방법을 적용하여 건설현장의 장비 및 자재의 상태와 상황을 정확하게 인식할 수 있는 개선된 모니터링 프레임워크를 개발하는 것을 목표로 한다. 첫 번째 연구 목표는 단일 카메라 기반의 깊이 인식 객체 탐지 프레임워크를 개발하여, 객체와 주변 표면 간 상대적 거리 등 공간적 단서를 활용해 객체의 상태를 인식하는 것이다. 건설 현장에서 크레인으로 자재를 들어 올리는 양중작업이 이러한 특성을 가지는 대표적인 사례이다. 일반적인 비전 모델은 매달린 자재와 지면에 놓인 자재가 이미지에서 매우 유사한 외형을 보일 때 이를 구분하는 데 어려움을 겪는다. 이러한 양중물을 식별하는 것은 현장 안전 확보를 위해 필수적이다. 그러나 외형 기반의 기존 탐지 모델은 매달린 객체와 그 아래 표면 간의 물리적 분리를 공간적으로 해석할 능력이 부족하다. 이를 극복하기 위해 본 연구에서는 깊이 추정과 정제된 공간적 특성 추출을 결합하여 양중물(Hanging Object)과 배경 사이의 차이 및 깊이의 불연속성 같은 매달림 관련 단서를 강조하는 프레임워크를 제안한다. 이를 지원하기 위해 실제 건설 현장 CCTV로부터 매달린 객체와 로프 등의 라벨이 포함된 101,381장의 이미지를 가진 HangCon 벤치마크 데이터셋을 구축했다. 또한, HODNet이라는 새로운 모듈을 개발하여 매달림과 관련된 깊이 특성을 추출하고, 전처리로 단일 카메라 깊이 추정의 정확도와 선명도를 개선하였다. 실험 결과, 제안된 깊이 인식 탐지 프레임워크는 특히 복잡하고 혼잡한 건설 환경에서 지면에 놓인 객체와 매달린 객체를 구분하는 데 있어 기존 방법 대비 정확도를 크게 향상했다. 두 번째 연구 목표는 터널 굴착과 같은 밀폐된 건설 환경에서 장비의 작동 상태를 해석하기 위한 맥락적 오디오-비주얼 다중 모달 인식 프레임워크를 구축하는 것이다. 이러한 환경은 제한된 시야, 빈번한 가림 현상, 여러 장비 간의 중첩된 활동 등으로 특징지어져 일반적인 비전 모델로 장비의 상태를 정확히 추론하기 어렵다. 장비의 실제 가동 여부를 인식하는 것은 작업 흐름 조정과 안전 관리에 중요하지만, 제한된 조건에서는 시각 정보만으로 충분한 단서를 얻기 어렵다. 이를 해결하기 위해 시각적 관찰과 청각 신호를 결합하고, 장비의 위치 및 작동 주기를 포착하는 공간적·시간적 맥락 모듈을 포함한 프레임워크를 제안한다. 특히 오디오 모달리티는 엔진 소리나 드릴링 소리와 같이 장비 고유의 음향 패턴을 탐지하는 데 중요한 역할을 수행하며, 작업 영역과의 공간적 근접성과 반복적인 작업 패턴을 통해 장비의 작동 상태를 효과적으로 추론할 수 있다. 이 프레임워크는 온라인 영상과 실제 터널 건설 현장 영상 데이터를 활용하여 검증되었으며, 단일 장비 및 다중 장비 상황에서 기존 비전 전용 시스템 대비 장비 활동 인식 성능을 크게 개선했다. 본 논문은 건설 장비와 자재의 상태를 인지할 수 있는 프레임워크를 제안하고 검증함으로써 건설 현장 모니터링 자동화 분야를 발전시켰다. 제안된 접근법은 시각적으로 모호하거나 복잡한 실제 환경에서도 확장 가능한 모니터링 시스템으로 확장할 수 있다. 이러한 기여는 외형 기반 및 단일 모달리티 모델의 주요 한계를 해결하며, 기존 인프라를 활용하여 지능형 비전 기반 모니터링 시스템을 실제 현장에 구축할 수 있는 실질적 기반을 마련한다. 더불어 공간적 정보와 소리 정보 등을 구조적으로 통합하는 방법론을 통하여, 장비 및 자재 인식에 새로운 접근을 제시하였고, 비전 기반 모니터링의 한계가 있는 다른 분야에서도 적용 가능한 일반화된 프레임워크를 제공한다.
    번역하기

    건설 프로젝트의 규모와 복잡성이 커짐에 따라 안전, 효율성, 상황 인식을 향상시킬 수 있는 자동화된 모니터링 시스템의 필요성이 커지고 있다. 최근 컴퓨터 비전 기반 모니터링 기술의 발...

    건설 프로젝트의 규모와 복잡성이 커짐에 따라 안전, 효율성, 상황 인식을 향상시킬 수 있는 자동화된 모니터링 시스템의 필요성이 커지고 있다. 최근 컴퓨터 비전 기반 모니터링 기술의 발전으로 작업자, 장비, 자재 등의 탐지 및 활동 인식을 가능하게 했지만, 기존 모델들은 시각적으로 모호하거나 문맥적으로 복잡한 조건에서 자재 및 장비의 작동 상태를 추론하는 데 한계가 있다. 본 논문에서는 비전 기반 모니터링 시스템의 두 가지 근본적인 한계를 다루고자 한다. 첫째, 기존 객체 탐지 모델은 형상, 색상, 질감과 같은 시각적 외형 단서에 크게 의존한다. 이러한 특징은 일반적인 객체 탐지에는 적합하지만, 특히 개별 상태가 유사한 외형을 보이는 경우 자재 및 장비의 작동 상태를 식별하는 데는 부적합하다. 이러한 한계는 건설 환경에서 문맥 의존적인 상황을 자주 오인하게 만들 수 있다. 둘째, 현재 접근 방식은 단일 모달리티 모델에 국한되어 있으며, 깊이 데이터, 소리, 시간적 패턴과 같은 보완적인 문맥 단서 없이 시각 정보만을 활용하는 경우가 많다. 이러한 비전 기반 모니터링 시스템은 가림 현상, 객체 간의 복잡한 상호작용이 많이 발생하는 동적인 건설 환경에서 문맥적 추론을 수행하는데 매우 제한적이다. 이러한 한계를 해결하기 위해 본 논문은 공간 추론 및 멀티 모달리티 방법을 적용하여 건설현장의 장비 및 자재의 상태와 상황을 정확하게 인식할 수 있는 개선된 모니터링 프레임워크를 개발하는 것을 목표로 한다. 첫 번째 연구 목표는 단일 카메라 기반의 깊이 인식 객체 탐지 프레임워크를 개발하여, 객체와 주변 표면 간 상대적 거리 등 공간적 단서를 활용해 객체의 상태를 인식하는 것이다. 건설 현장에서 크레인으로 자재를 들어 올리는 양중작업이 이러한 특성을 가지는 대표적인 사례이다. 일반적인 비전 모델은 매달린 자재와 지면에 놓인 자재가 이미지에서 매우 유사한 외형을 보일 때 이를 구분하는 데 어려움을 겪는다. 이러한 양중물을 식별하는 것은 현장 안전 확보를 위해 필수적이다. 그러나 외형 기반의 기존 탐지 모델은 매달린 객체와 그 아래 표면 간의 물리적 분리를 공간적으로 해석할 능력이 부족하다. 이를 극복하기 위해 본 연구에서는 깊이 추정과 정제된 공간적 특성 추출을 결합하여 양중물(Hanging Object)과 배경 사이의 차이 및 깊이의 불연속성 같은 매달림 관련 단서를 강조하는 프레임워크를 제안한다. 이를 지원하기 위해 실제 건설 현장 CCTV로부터 매달린 객체와 로프 등의 라벨이 포함된 101,381장의 이미지를 가진 HangCon 벤치마크 데이터셋을 구축했다. 또한, HODNet이라는 새로운 모듈을 개발하여 매달림과 관련된 깊이 특성을 추출하고, 전처리로 단일 카메라 깊이 추정의 정확도와 선명도를 개선하였다. 실험 결과, 제안된 깊이 인식 탐지 프레임워크는 특히 복잡하고 혼잡한 건설 환경에서 지면에 놓인 객체와 매달린 객체를 구분하는 데 있어 기존 방법 대비 정확도를 크게 향상했다. 두 번째 연구 목표는 터널 굴착과 같은 밀폐된 건설 환경에서 장비의 작동 상태를 해석하기 위한 맥락적 오디오-비주얼 다중 모달 인식 프레임워크를 구축하는 것이다. 이러한 환경은 제한된 시야, 빈번한 가림 현상, 여러 장비 간의 중첩된 활동 등으로 특징지어져 일반적인 비전 모델로 장비의 상태를 정확히 추론하기 어렵다. 장비의 실제 가동 여부를 인식하는 것은 작업 흐름 조정과 안전 관리에 중요하지만, 제한된 조건에서는 시각 정보만으로 충분한 단서를 얻기 어렵다. 이를 해결하기 위해 시각적 관찰과 청각 신호를 결합하고, 장비의 위치 및 작동 주기를 포착하는 공간적·시간적 맥락 모듈을 포함한 프레임워크를 제안한다. 특히 오디오 모달리티는 엔진 소리나 드릴링 소리와 같이 장비 고유의 음향 패턴을 탐지하는 데 중요한 역할을 수행하며, 작업 영역과의 공간적 근접성과 반복적인 작업 패턴을 통해 장비의 작동 상태를 효과적으로 추론할 수 있다. 이 프레임워크는 온라인 영상과 실제 터널 건설 현장 영상 데이터를 활용하여 검증되었으며, 단일 장비 및 다중 장비 상황에서 기존 비전 전용 시스템 대비 장비 활동 인식 성능을 크게 개선했다. 본 논문은 건설 장비와 자재의 상태를 인지할 수 있는 프레임워크를 제안하고 검증함으로써 건설 현장 모니터링 자동화 분야를 발전시켰다. 제안된 접근법은 시각적으로 모호하거나 복잡한 실제 환경에서도 확장 가능한 모니터링 시스템으로 확장할 수 있다. 이러한 기여는 외형 기반 및 단일 모달리티 모델의 주요 한계를 해결하며, 기존 인프라를 활용하여 지능형 비전 기반 모니터링 시스템을 실제 현장에 구축할 수 있는 실질적 기반을 마련한다. 더불어 공간적 정보와 소리 정보 등을 구조적으로 통합하는 방법론을 통하여, 장비 및 자재 인식에 새로운 접근을 제시하였고, 비전 기반 모니터링의 한계가 있는 다른 분야에서도 적용 가능한 일반화된 프레임워크를 제공한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    As the complexity and scale of construction projects continue to increase, the demand for intelligent monitoring systems that can enhance safety, efficiency, and situational awareness has become increasingly critical. While recent advances in vision-based monitoring has enabled object detection and activity recognition, conventional models remain limited in their ability to infer operational states of materials and equipment, especially under visually ambiguous or contextually complex conditions. This dissertation addresses two fundamental limitations in current vision-based monitoring systems. First, conventional object detection models rely heavily on visual appearance cues such as shape, color, and texture. While these features are sufficient for basic object recognition, they are inadequate for identifying the operational states of materials and equipment due to visual ambiguity when n different states exhibit similar external appearances. This challenge leads to frequent misinterpretation of context-dependent conditions in construction environments. Second, current approaches are limited to a single sensing modality, typically relying solely on visual data without incorporating complementary contextual cues such as depth, audio signals, or temporal patterns. This reliance on a single modality significantly restricts the capacity of monitoring systems to perform contextual reasoning, particularly in dynamic construction environments characterized by occlusion, clutter, and complex interactions among entities. To address these challenges, this dissertation aims to develop an advanced monitoring framework for construction environments that enables accurate, context-aware recognition of material and equipment operational states by integrating spatial and multimodal information. The first objective focuses on developing a monocular depth-aware object detection framework that enables recognition of object states by incorporating spatial cues derived from depth information, such as the relative distance between objects and their surrounding surfaces. The recognition of hanging objects during lifting operations in construction site is a representative case of this challenge. In this scenario, conventional vision-based models often struggle to differentiate between materials that are hanging and those placed on the ground, as both can exhibit highly similar visual appearances in RGB images. Identifying whether an object is in a hanging state is critical for assessing site safety, particularly in relation to crane operations and material handling. However, appearance-based detectors lack the spatial understanding required to interpret the physical disconnection between a hanging object and the surface below. To address this limitation, the proposed framework integrates depth estimation and refined spatial feature extraction to highlight suspension-related cues, such as vertical gaps and depth discontinuities. To support this framework, a benchmark dataset named HangCon was constructed, consisting of 101,381 images annotated with context-aware labels including hanging object and rope, enabling robust detection of suspended materials in real-world scenes. Furthermore, a novel module called HODNet was introduced to extract suspension-relevant depth features, utilizing segmentation-guided preprocessing to improve the clarity and accuracy of monocular depth estimates. Experimental results demonstrated that the proposed depth-aware detection framework achieves significantly improved accuracy in distinguishing hanging from grounded objects, especially in cluttered and visually complex construction environments. The second objective centers on establishing a contextual audio-visual multimodal recognition framework for interpreting equipment operations in closed-space construction environments, such as tunnel excavation sites. These environments are typically characterized by limited visibility, frequent occlusion, and overlapping activities among multiple equipment, making it difficult for conventional vision-only models to reliably infer equipment states. Recognizing whether equipment is actively operating or idle is critical for workflow coordination and safety management, but visual appearance alone often fails to provide sufficient cues especially under conditions of poor lighting, dust, or restricted viewpoints. To overcome these challenges, the proposed framework integrates auditory signals with visual observations, along with spatial and temporal context modules that capture the positioning and operational cycles of equipment. The audio modality plays a particularly important role in detecting equipment-specific acoustic patterns, such as engine noise or drilling sounds, which often correlate with functional status. Spatial proximity to key working areas is used to infer likely operational engagement, while cyclical temporal modeling captures repetitive patterns of activity typical in tunneling workflows. The framework was evaluated on both curated online videos and real-world surveillance footage collected from tunnel construction sites, encompassing single-equipment and multi-equipment scenarios. Results demonstrate that the integration of multimodal and contextual cues significantly improves the model’s ability to recognize concurrent equipment activities compared to baseline vision-only systems, especially in visually constrained or acoustically complex conditions. This dissertation advances the field of automated construction monitoring by demonstrating that context-aware recognition of operational states can be achieved without reliance on specialized sensors or domain-specific heuristics. Through the use of monocular depth estimation and ambient audio signals, the proposed approaches enable scalable and non-intrusive monitoring across a range of real-world site conditions, including visually ambiguous or acoustically complex environments. These contributions address key limitations of appearance-based and single-modality models, and establish a practical foundation for deploying intelligent monitoring systems using existing infrastructure. Moreover, the dissertation offers methodological innovations by formalizing structured techniques for incorporating spatial disconnection and temporal regularity into operational state recognition, providing a generalizable framework applicable to other domains where visual cues alone are insufficient.
    번역하기

    As the complexity and scale of construction projects continue to increase, the demand for intelligent monitoring systems that can enhance safety, efficiency, and situational awareness has become increasingly critical. While recent advances in vision-b...

    As the complexity and scale of construction projects continue to increase, the demand for intelligent monitoring systems that can enhance safety, efficiency, and situational awareness has become increasingly critical. While recent advances in vision-based monitoring has enabled object detection and activity recognition, conventional models remain limited in their ability to infer operational states of materials and equipment, especially under visually ambiguous or contextually complex conditions. This dissertation addresses two fundamental limitations in current vision-based monitoring systems. First, conventional object detection models rely heavily on visual appearance cues such as shape, color, and texture. While these features are sufficient for basic object recognition, they are inadequate for identifying the operational states of materials and equipment due to visual ambiguity when n different states exhibit similar external appearances. This challenge leads to frequent misinterpretation of context-dependent conditions in construction environments. Second, current approaches are limited to a single sensing modality, typically relying solely on visual data without incorporating complementary contextual cues such as depth, audio signals, or temporal patterns. This reliance on a single modality significantly restricts the capacity of monitoring systems to perform contextual reasoning, particularly in dynamic construction environments characterized by occlusion, clutter, and complex interactions among entities. To address these challenges, this dissertation aims to develop an advanced monitoring framework for construction environments that enables accurate, context-aware recognition of material and equipment operational states by integrating spatial and multimodal information. The first objective focuses on developing a monocular depth-aware object detection framework that enables recognition of object states by incorporating spatial cues derived from depth information, such as the relative distance between objects and their surrounding surfaces. The recognition of hanging objects during lifting operations in construction site is a representative case of this challenge. In this scenario, conventional vision-based models often struggle to differentiate between materials that are hanging and those placed on the ground, as both can exhibit highly similar visual appearances in RGB images. Identifying whether an object is in a hanging state is critical for assessing site safety, particularly in relation to crane operations and material handling. However, appearance-based detectors lack the spatial understanding required to interpret the physical disconnection between a hanging object and the surface below. To address this limitation, the proposed framework integrates depth estimation and refined spatial feature extraction to highlight suspension-related cues, such as vertical gaps and depth discontinuities. To support this framework, a benchmark dataset named HangCon was constructed, consisting of 101,381 images annotated with context-aware labels including hanging object and rope, enabling robust detection of suspended materials in real-world scenes. Furthermore, a novel module called HODNet was introduced to extract suspension-relevant depth features, utilizing segmentation-guided preprocessing to improve the clarity and accuracy of monocular depth estimates. Experimental results demonstrated that the proposed depth-aware detection framework achieves significantly improved accuracy in distinguishing hanging from grounded objects, especially in cluttered and visually complex construction environments. The second objective centers on establishing a contextual audio-visual multimodal recognition framework for interpreting equipment operations in closed-space construction environments, such as tunnel excavation sites. These environments are typically characterized by limited visibility, frequent occlusion, and overlapping activities among multiple equipment, making it difficult for conventional vision-only models to reliably infer equipment states. Recognizing whether equipment is actively operating or idle is critical for workflow coordination and safety management, but visual appearance alone often fails to provide sufficient cues especially under conditions of poor lighting, dust, or restricted viewpoints. To overcome these challenges, the proposed framework integrates auditory signals with visual observations, along with spatial and temporal context modules that capture the positioning and operational cycles of equipment. The audio modality plays a particularly important role in detecting equipment-specific acoustic patterns, such as engine noise or drilling sounds, which often correlate with functional status. Spatial proximity to key working areas is used to infer likely operational engagement, while cyclical temporal modeling captures repetitive patterns of activity typical in tunneling workflows. The framework was evaluated on both curated online videos and real-world surveillance footage collected from tunnel construction sites, encompassing single-equipment and multi-equipment scenarios. Results demonstrate that the integration of multimodal and contextual cues significantly improves the model’s ability to recognize concurrent equipment activities compared to baseline vision-only systems, especially in visually constrained or acoustically complex conditions. This dissertation advances the field of automated construction monitoring by demonstrating that context-aware recognition of operational states can be achieved without reliance on specialized sensors or domain-specific heuristics. Through the use of monocular depth estimation and ambient audio signals, the proposed approaches enable scalable and non-intrusive monitoring across a range of real-world site conditions, including visually ambiguous or acoustically complex environments. These contributions address key limitations of appearance-based and single-modality models, and establish a practical foundation for deploying intelligent monitoring systems using existing infrastructure. Moreover, the dissertation offers methodological innovations by formalizing structured techniques for incorporating spatial disconnection and temporal regularity into operational state recognition, providing a generalizable framework applicable to other domains where visual cues alone are insufficient.

    더보기

    목차 (Table of Contents)

    • Chapter 1. Introduction 1
    • 1.1 Research Background 1
    • 1.2 Problem Description 4
    • 1.3 Research Objectives and Scope 8
    • 1.4 Dissertation Outline 18
    • Chapter 1. Introduction 1
    • 1.1 Research Background 1
    • 1.2 Problem Description 4
    • 1.3 Research Objectives and Scope 8
    • 1.4 Dissertation Outline 18
    • Chapter 2. Theoretical Backgrounds 20
    • 2.1 Vision-Based Monitoring in Construction 20
    • 2.1.1 Object Detection Model for Construction 21
    • 2.1.2 Applications in Equipment and Worker Monitoring 23
    • 2.1.3 Applications in Hanging Object Monitoring 25
    • 2.1.4 Challenges of Object Detection in Construction Monitoring 27
    • 2.2 Multimodal and Context-Aware Monitoring 29
    • 2.2.1 Audio-Based Monitoring in Construction 29
    • 2.2.2 Multimodal Approaches for Monitoring 32
    • 2.2.3 Context-Aware and Sequential Recognition 35
    • 2.3 Construction-Specific Dataset 38
    • 2.3.1 Image Dataset for Vision-Based Monitoring 39
    • 2.3.2 Multimodal Dataset for Construction Monitoring 41
    • 2.4 Summary 42
    • Chapter 3. Methodology 43
    • 3.1 Monocular Depth-Aware Object Detection Model 45
    • 3.1.1 Motivation and Design Principles 45
    • 3.1.2 Overall Architecture 47
    • 3.1.3 Module Design 49
    • 3.1.4 Integration Strategy 53
    • 3.2 Audio-Visual Multimodal Recognition Framework 55
    • 3.2.1 Motivation 55
    • 3.2.2 Overall Framework 57
    • 3.2.3 Visual Feature Extraction 60
    • 3.2.4 Audio Feature Extraction 62
    • 3.2.5 Feature Fusion and Classification 64
    • Chapter 4. HangCon: Benchmark Dataset for Enhanced Detection of Hanging Objects in Construction Sites 66
    • 4.1 Motivation and Problem Formulation 67
    • 4.2 HangCon: Benchmark Dataset for Hanging Object 73
    • 4.2.1 Data Collection 75
    • 4.2.2 Data Preprocessing 80
    • 4.2.3 Data Annotation 81
    • 4.2.4 Data Split 85
    • 4.3 Benchmark Design 87
    • 4.3.1 Evaluation Protocol for Object Detection 87
    • 4.3.2 Evaluation Protocol for Image Classification 94
    • 4.4 Experiment Results and Discussions 96
    • 4.4.1 Results of Object Detection 96
    • 4.4.2 Results of Image Classification 104
    • 4.4.3 Discussion 108
    • 4.5 Summary 115
    • Chapter 5. Monocular Depth-Aware Model for Detecting Hanging Objects 117
    • 5.1 Motivation and Overall Framework 119
    • 5.2 Segmentation-Guided Depth Preprocessing 124
    • 5.2.1 Monocular Depth Map Extraction 126
    • 5.2.2 Segmentation Map Extraction 127
    • 5.2.3 Edge-preserving Smoothing 129
    • 5.2.4 Depth Boundary Enhancement 130
    • 5.3 Depth-aware Hanging Object Detection Model 132
    • 5.4 Experiment Design 135
    • 5.4.1 Dataset 135
    • 5.4.2 Evaluation Metrics 138
    • 5.4.3 Computing Models 140
    • 5.5 Results and Discussion 141
    • 5.5.1 Evaluation of Depth-Aware Modeling 141
    • 5.5.2 Effect of Rope Annotations 146
    • 5.5.3 Generalization to Out-of-Distribution Images 149
    • 5.6 VLM for Hanging Object Detection 153
    • 5.6.1 Models for Comparison 153
    • 5.6.2 Experiment Settings 156
    • 5.6.3 Experimental Results and Discussion 159
    • 5.7 Summary 170
    • Chapter 6. Contextual Audio-Visual Multimodal Model for Recognizing Concurrent Activities of Equipment 173
    • 6.1 Motivation 175
    • 6.2 Development of Audio-Visual Multimodal Model for Single-Equipment Action Recognition 178
    • 6.3 Development of Contextual Audio-Visual Multimodal Model for Multi-Equipment Activity Recognition 181
    • 6.3.1 Vision Module 184
    • 6.3.2 Audio Module 188
    • 6.3.3 Spatial context module 189
    • 6.3.4 Cyclical temporal context module 192
    • 6.3.5 Integration of modules 194
    • 6.4 Experiment Design 195
    • 6.4.1 Datasets 195
    • 6.4.2 Audio-Visual Action Recognition for Single Equipment 202
    • 6.4.3 Contextual Audio-Visual Activity Recognition for Multi-Equipment Operations 203
    • 6.4.4 Implementation Details 205
    • 6.5 Results and Discussion 206
    • 6.5.1 Performance on Single-Equipment Action Classification 206
    • 6.5.2 Performance on Multi-Equipment Activity Classification 211
    • 6.6 Summary 219
    • Chapter 7. Conclusions 221
    • 7.1 Achievements to Research Objectives 221
    • 7.2 Contributions 225
    • 7.3 Future Research 227
    • References 229
    • 국 문 초 록 240
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼