RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Context-Aware Vision–Language Safety Reasoning for UxVs = Context-Aware Vision–Language Safety Reasoning for UxVs

    한글로보기

    https://www.riss.kr/link?id=T17385649

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This thesis addresses real-time safety reasoning for unmanned ground and aerial vehicles (UxVs) operating on resource-constrained edge platforms. Practical field operation requires not only robust perception in challenging scenes (low light, clutter, occlusion) but also human-understandable explanations that support operator trust and rapid decision making. To meet these needs, this thesis presents SafeVision, a parameter-efficient vision language system that delivers scene-level understanding, region-grounded safety checks, and visual question answering. SafeVision couples a froze CLIP-ViT visual encoder with two complementary reasoning paths: a Scene-Aware Reasoner (SAR) that captures global context and a Region Focused Neural Reasoner (ReFiNER) that localizes PPE/hazard evidence. A lightweight cross-modal adapter and task tokens ([scene], [region], [vqa]) condition a LoRA-tuned language model to produce grounded, interpretable outputs while keeping trainable parameters small for edge deployment. The approach is evaluated on a custom industrial-safety dataset with scene, region, and VQA annotations. Empirically, SAR achieves 87.4% scene accuracy, and the VQA head attains BLEU-4 83.2, F1 84.6, EM 80.3, and METEOR 82.5, with qualitative examples demonstrating clear alignment between visual evidence and textual reasoning. Prototype deployments on UAV/UGV platforms show that SafeVision meets typical on-board latency and memory budgets.
    번역하기

    This thesis addresses real-time safety reasoning for unmanned ground and aerial vehicles (UxVs) operating on resource-constrained edge platforms. Practical field operation requires not only robust perception in challenging scenes (low light, clutter, ...

    This thesis addresses real-time safety reasoning for unmanned ground and aerial vehicles (UxVs) operating on resource-constrained edge platforms. Practical field operation requires not only robust perception in challenging scenes (low light, clutter, occlusion) but also human-understandable explanations that support operator trust and rapid decision making. To meet these needs, this thesis presents SafeVision, a parameter-efficient vision language system that delivers scene-level understanding, region-grounded safety checks, and visual question answering. SafeVision couples a froze CLIP-ViT visual encoder with two complementary reasoning paths: a Scene-Aware Reasoner (SAR) that captures global context and a Region Focused Neural Reasoner (ReFiNER) that localizes PPE/hazard evidence. A lightweight cross-modal adapter and task tokens ([scene], [region], [vqa]) condition a LoRA-tuned language model to produce grounded, interpretable outputs while keeping trainable parameters small for edge deployment. The approach is evaluated on a custom industrial-safety dataset with scene, region, and VQA annotations. Empirically, SAR achieves 87.4% scene accuracy, and the VQA head attains BLEU-4 83.2, F1 84.6, EM 80.3, and METEOR 82.5, with qualitative examples demonstrating clear alignment between visual evidence and textual reasoning. Prototype deployments on UAV/UGV platforms show that SafeVision meets typical on-board latency and memory budgets.

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼