As construction sites become increasingly complex and risk-prone, the demand for intelligent safety monitoring systems capable of recognizing both visible hazards and latent risk factors continues to grow. This study presents a Safety-Aware Image Capt...
As construction sites become increasingly complex and risk-prone, the demand for intelligent safety monitoring systems capable of recognizing both visible hazards and latent risk factors continues to grow. This study presents a Safety-Aware Image Captioning framework that adapts large vision-language models (VLMs) to the domain of construction safety through structured fine-tuning. In contrast to general-purpose captioning systems trained on open-domain datasets, the proposed approach incorporates regulatory metadata and accident scenario taxonomies grounded in OSHA and KOSHA standards to enable contextual hazard reasoning.
A synthetic dataset comprising 3,000 images was developed, covering ten accident item categories, with each image annotated using structured safety metadata—such as PPE compliance, fall protection status, and hazard-specific visual cues. The InternVL 2.5 8B model was fine-tuned under structured and unstructured training regimes and compared to strong baselines, including GPT-4o and a zero-shot InternVL configuration. A custom test set of 50 real-world construction images was curated and manually annotated with expert reference captions to support robust evaluation.
Quantitative results revealed that the structured-trained InternVL model achieved the highest scores across all metrics, with a BLEU-4 of 0.384 and a CIDEr of 1.79—marking a substantial improvement in semantic alignment and regulatory interpretability over all baselines. Qualitative analysis further demonstrated that this model consistently identified compliance violations and inferred latent hazards, including unanchored ladders, missing guardrails, excavation fall risks, and heavy equipment proximity threats.
These findings confirm that structured domain adaptation substantially improves hazard reasoning capabilities in vision-language models. Despite limitations in environmental realism and temporal context, the proposed framework provides a scalable, interpretable, and reproducible method for automated safety monitoring in high-risk industrial settings. As the current system is based on static imagery, future research should explore the integration of video-based temporal reasoning (e.g., video-to-text captioning) to support real-time hazard detection and prediction.