RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Safety Hazard Interpretation using Diffusion-generated Construction Images and Vision-Language Models Trained by Accident Cases = 건설 사고사례로 학습된 고품질 생성형 확산 모델과 비전-언어 모델을 활용한 현장 안전 위험성 분석 기술 개발

    한글로보기

    https://www.riss.kr/link?id=T17315168

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    As construction sites become increasingly complex and risk-prone, the demand for intelligent safety monitoring systems capable of recognizing both visible hazards and latent risk factors continues to grow. This study presents a Safety-Aware Image Captioning framework that adapts large vision-language models (VLMs) to the domain of construction safety through structured fine-tuning. In contrast to general-purpose captioning systems trained on open-domain datasets, the proposed approach incorporates regulatory metadata and accident scenario taxonomies grounded in OSHA and KOSHA standards to enable contextual hazard reasoning.
    A synthetic dataset comprising 3,000 images was developed, covering ten accident item categories, with each image annotated using structured safety metadata—such as PPE compliance, fall protection status, and hazard-specific visual cues. The InternVL 2.5 8B model was fine-tuned under structured and unstructured training regimes and compared to strong baselines, including GPT-4o and a zero-shot InternVL configuration. A custom test set of 50 real-world construction images was curated and manually annotated with expert reference captions to support robust evaluation.
    Quantitative results revealed that the structured-trained InternVL model achieved the highest scores across all metrics, with a BLEU-4 of 0.384 and a CIDEr of 1.79—marking a substantial improvement in semantic alignment and regulatory interpretability over all baselines. Qualitative analysis further demonstrated that this model consistently identified compliance violations and inferred latent hazards, including unanchored ladders, missing guardrails, excavation fall risks, and heavy equipment proximity threats.
    These findings confirm that structured domain adaptation substantially improves hazard reasoning capabilities in vision-language models. Despite limitations in environmental realism and temporal context, the proposed framework provides a scalable, interpretable, and reproducible method for automated safety monitoring in high-risk industrial settings. As the current system is based on static imagery, future research should explore the integration of video-based temporal reasoning (e.g., video-to-text captioning) to support real-time hazard detection and prediction.
    번역하기

    As construction sites become increasingly complex and risk-prone, the demand for intelligent safety monitoring systems capable of recognizing both visible hazards and latent risk factors continues to grow. This study presents a Safety-Aware Image Capt...

    As construction sites become increasingly complex and risk-prone, the demand for intelligent safety monitoring systems capable of recognizing both visible hazards and latent risk factors continues to grow. This study presents a Safety-Aware Image Captioning framework that adapts large vision-language models (VLMs) to the domain of construction safety through structured fine-tuning. In contrast to general-purpose captioning systems trained on open-domain datasets, the proposed approach incorporates regulatory metadata and accident scenario taxonomies grounded in OSHA and KOSHA standards to enable contextual hazard reasoning.
    A synthetic dataset comprising 3,000 images was developed, covering ten accident item categories, with each image annotated using structured safety metadata—such as PPE compliance, fall protection status, and hazard-specific visual cues. The InternVL 2.5 8B model was fine-tuned under structured and unstructured training regimes and compared to strong baselines, including GPT-4o and a zero-shot InternVL configuration. A custom test set of 50 real-world construction images was curated and manually annotated with expert reference captions to support robust evaluation.
    Quantitative results revealed that the structured-trained InternVL model achieved the highest scores across all metrics, with a BLEU-4 of 0.384 and a CIDEr of 1.79—marking a substantial improvement in semantic alignment and regulatory interpretability over all baselines. Qualitative analysis further demonstrated that this model consistently identified compliance violations and inferred latent hazards, including unanchored ladders, missing guardrails, excavation fall risks, and heavy equipment proximity threats.
    These findings confirm that structured domain adaptation substantially improves hazard reasoning capabilities in vision-language models. Despite limitations in environmental realism and temporal context, the proposed framework provides a scalable, interpretable, and reproducible method for automated safety monitoring in high-risk industrial settings. As the current system is based on static imagery, future research should explore the integration of video-based temporal reasoning (e.g., video-to-text captioning) to support real-time hazard detection and prediction.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    복잡성과 위험성이 증가하는 건설현장에서 시각적으로 명확한 위험뿐 아니라 잠재적인 위험요소까지 인식 가능한 지능형 안전 모니터링 시스템에 대한 요구가 높아지고 있다. 본 연구는 범용 시각-언어 모델(Vision-Language Model)을 건설안전 도메인에 특화된 방식으로 파인튜닝하여 적용하는 Safety-Aware 이미지 캡셔닝 프레임워크를 제안한다. 제안된 접근법은 기존 개방형 데이터셋 기반 모델과 달리, OSHA 및 KOSHA 기준에 기반한 사고 시나리오 분류 및 안전 규정 메타데이터를 통합함으로써 상황 맥락 기반의 위험 추론을 가능하게 한다.
    총 3,000장의 합성 이미지가 10개 사고 유형 범주를 기준으로 생성되었으며, 각 이미지에는 PPE 착용 여부, 낙상 방지 장치 유무, 시나리오별 위험요소 등을 포함한 구조화된 메타데이터 기반 캡션이 주석화되었다. InternVL 2.5 8B 모델은 구조화 및 비구조화 학습 방식으로 각각 파인튜닝되었으며, GPT-4o 및 Zero-shot InternVL과의 성능 비교를 통해 평가되었다. 또한, 실제 건설현장에서 수집한 50장의 테스트 이미지를 기반으로, 연구자가 직접 다수의 참조 캡션을 생성하여 정량적·정성적 분석을 수행하였다.
    평가 결과, 구조화 학습된 InternVL 모델은 BLEU-4 0.384, CIDEr 1.79의 최고 성능을 기록하며 모든 비교 모델을 상회하였다. 정성적 비교에서도, 본 모델은 사다리 고정 미비, 작업대 가드레일 부재, 굴착 작업 중 낙상 위험, 장비 접근 위험 등 실제 사고 가능성과 관련된 복합적 위험 요소들을 일관되게 식별하고 설명하였다.
    이러한 결과는 구조화된 안전 도메인 지식을 활용한 파인튜닝이 멀티모달 모델의 위험 추론 능력을 효과적으로 향상시킴을 입증한다. 다만, 환경적 현실감 부족 및 시간적 정보 미반영 등 일부 한계가 존재하며, 향후 연구에서는 영상 기반의 시계열 추론(video-to-text)을 통해 실시간 위험 탐지와 예측 기능을 통합하는 방향으로 확장할 필요가 있다.
    번역하기

    복잡성과 위험성이 증가하는 건설현장에서 시각적으로 명확한 위험뿐 아니라 잠재적인 위험요소까지 인식 가능한 지능형 안전 모니터링 시스템에 대한 요구가 높아지고 있다. 본 연구는 범...

    복잡성과 위험성이 증가하는 건설현장에서 시각적으로 명확한 위험뿐 아니라 잠재적인 위험요소까지 인식 가능한 지능형 안전 모니터링 시스템에 대한 요구가 높아지고 있다. 본 연구는 범용 시각-언어 모델(Vision-Language Model)을 건설안전 도메인에 특화된 방식으로 파인튜닝하여 적용하는 Safety-Aware 이미지 캡셔닝 프레임워크를 제안한다. 제안된 접근법은 기존 개방형 데이터셋 기반 모델과 달리, OSHA 및 KOSHA 기준에 기반한 사고 시나리오 분류 및 안전 규정 메타데이터를 통합함으로써 상황 맥락 기반의 위험 추론을 가능하게 한다.
    총 3,000장의 합성 이미지가 10개 사고 유형 범주를 기준으로 생성되었으며, 각 이미지에는 PPE 착용 여부, 낙상 방지 장치 유무, 시나리오별 위험요소 등을 포함한 구조화된 메타데이터 기반 캡션이 주석화되었다. InternVL 2.5 8B 모델은 구조화 및 비구조화 학습 방식으로 각각 파인튜닝되었으며, GPT-4o 및 Zero-shot InternVL과의 성능 비교를 통해 평가되었다. 또한, 실제 건설현장에서 수집한 50장의 테스트 이미지를 기반으로, 연구자가 직접 다수의 참조 캡션을 생성하여 정량적·정성적 분석을 수행하였다.
    평가 결과, 구조화 학습된 InternVL 모델은 BLEU-4 0.384, CIDEr 1.79의 최고 성능을 기록하며 모든 비교 모델을 상회하였다. 정성적 비교에서도, 본 모델은 사다리 고정 미비, 작업대 가드레일 부재, 굴착 작업 중 낙상 위험, 장비 접근 위험 등 실제 사고 가능성과 관련된 복합적 위험 요소들을 일관되게 식별하고 설명하였다.
    이러한 결과는 구조화된 안전 도메인 지식을 활용한 파인튜닝이 멀티모달 모델의 위험 추론 능력을 효과적으로 향상시킴을 입증한다. 다만, 환경적 현실감 부족 및 시간적 정보 미반영 등 일부 한계가 존재하며, 향후 연구에서는 영상 기반의 시계열 추론(video-to-text)을 통해 실시간 위험 탐지와 예측 기능을 통합하는 방향으로 확장할 필요가 있다.

    더보기

    목차 (Table of Contents)

    • DEDICATION iii
    • ACKNOWLEDGEMENT iv
    • ABSTRACT vii
    • Contents x
    • List of Table xiii
    • DEDICATION iii
    • ACKNOWLEDGEMENT iv
    • ABSTRACT vii
    • Contents x
    • List of Table xiii
    • List of Figure xiv
    • Chapter 1. Introduction 1
    • 1.1. Research Background 1
    • 1.2. Problem Statement 8
    • 1.3. Research Objectives, Scope and Contributions 11
    • 1.4. Dissertation Outline 15
    • Chapter 2. Literature Review 18
    • 2.1. Vision-based Monitoring Technologies 19
    • 2.1.1. Object-Centric Safety Monitoring Approaches 19
    • 2.1.2. Contextual and Vision-Language-Based Monitoring Approaches 21
    • 2.2. Domain Aware Synthetic Image Generation for Safety Applications 24
    • 2.2.1. Traditional and Generative Approaches for Synthetic Image Generation 24
    • 2.2.2. Domain Adaptation Techniques for Safety-Focused Image Generation 26
    • 2.3. Limitations of Existing Studies 29
    • Chapter 3. Safety Scenarios Classification and Structuring 31
    • 3.1. Systematic Modeling of Accident Scenario 32
    • 3.1.1. Theoretical Foundations of Accident Scenario Analysis 32
    • 3.1.2. Development and Description of the Integrated Accident Scenario Model 36
    • 3.2. Statistical Analysis of Construction Accident Patterns 41
    • 3.2.1 Data Collection and Multi-Dimensional Structuring 41
    • 3.2.2 Intermediate Structuring and Preparation for Pattern Consolidation 51
    • 3.3. Development of Representative Safety Scenarios 56
    • Chapter 4. Hazard-Specific Text-to-Image Generation 63
    • 4.1. Selective Acquisition of Data and Preprocessing 65
    • 4.2. Implementation of LoRA-Based Fine-Tuning Framework 71
    • 4.3. Development of Hazard-Specific Image Generation Pipeline 78
    • 4.3.1. Multi-Modal Prompt Engineering and LoRA Integration 79
    • 4.3.2. Optimization of Sampling Parameters 85
    • 4.3.3. Quality Control and Selection of Generated Images 87
    • 4.4. Verification of Synthetic Data Effectiveness 91
    • Chapter 5. Safety Aware Image Captioning 93
    • 5.1. Additional Metadata Annotation for Domain Adaptation 95
    • 5.2. Domain-Specific Adaptation of a Vision-Language Model 100
    • 5.3. Implementation of Image Captioning System 102
    • Chapter 6. Experimental Results and Discussions 106
    • 6.1. Evaluation of Hazard-Specific Synthetic Image Quality 107
    • 6.2 Quantitative Evaluation of Captioning Model Performance 118
    • 6.3 Qualitative Analysis of Caption Generation 121
    • 6.4 Discussion 128
    • Chapter 7. Conclusions 133
    • 7.1. Summary and Contributions 134
    • 7.2. Improvement Opportunities and Future Research 137
    • Bibliography 140
    • 국문 초록 150
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼