RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Data-driven Anomaly Detection using Hierarchical Multi-labeled Classification for Vehicular Communications

    한글로보기

    https://www.riss.kr/link?id=T15784719

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    융합 기술의 발전으로 이기종 산업군의 시스템과 데이터는 네트워크로 상호 연결되고 있습니다. 커넥티비티 속성을 갖는 자율주행 자동차는 지능형 교통체계와 연계하여 도로 위의 혼잡지역을 실시간으로 회피하고, 사람과 가장 가까운 거리에서 24시간 이동통신망에 연결되는 스마트 기기와 웨어러블 기기는 심박 수, 활동량, 신체리듬 등의 건강 데이터를 수집하고 분석하여 다양한 형태의 헬스케어 서비스를 제공합니다. 이러한 기기간 융합은 사용자에게 편리함을 주지만, 동시에 해커에게는 공격의 접점을 제공할 수 있습니다. 날로 지능화되고 있는 사이버 공격을 사전에 원천적으로 차단하는 것은 사실상 불가능하기 때문에 오히려 빠른 탐지로 신속하게 조치방안을 찾는 것이 더 중요해지고 있습니다. 실시간 속성이 중요한 차량의 침입탐지에서는 제한된 시간 안에 오탐이나 탐지누락을 최소화하여 유입되는 위협을 식별하고 차단하는 것이 중요합니다. 이를 위해 본 논문은 차량의 통신에서 이상 행위를 탐지하는 모델의 성능 향상을 위해 머신러닝을 사용한 세 가지 주요 분석 모델을 제안합니다.

    •우리는 데이터셋의 전처리 단계에서 모델의 입력으로 사용하기 위한 필수 특징을 추출하기 위해 correlation-entropy feature selecion (CEFS) 방법을 제안합니다. 모델은 데이터 연관성과 가중치를 기반으로 추출한 특징을 활용하여 데이터 샘플을 학습하고, 그것이 정상인지 공격인지 판별합니다. 또한 하이퍼파라미터를 조정하는 반복된 학습을 통해 탐지 성능이 향상된 모델을 도출할 수 있습니다.
    •본 논문은 데이터의 대부분이 정상 샘플을 차지하는 불균형한 양상을 띄는 트래픽에서 극소수의 공격 데이터를 놓치지 않고 탐지하기 위해 오버샘플링과 언더샘플링을 조합한 샘플링 알고리즘 combined oversampling and undersampling based on slow-start (COUSS)를 제안합니다. 그 결과 적은 수를 갖는 공격 데이터가 탐지에서 누락되지 않고, 기존보다 높은 정확도를 도출하는 것을 확인할 수 있었습니다.
    •이 논문은 공격 샘플의 상세한 속성을 분류하기 위해 다중 레이블과 다중 클래스의 구조를 사용하는 계층적 분류 알고리즘을 제안합니다. 기존의 탐지 모델이 일반적으로 대상 데이터에 공격의 발생여부 즉, 공격인지 정상인지만 분류한다면, 본 논문의 방법은 공격의 발생여부 뿐만 아니라, 공격의 종류와 해당 공격이 발생한 차량의 종류, 추가적인 공격 속성 등 세부 공격 레이블을 빠르게 식별할 수 있음을 의미합니다. 본 모델의 기여는 대부분의 정상 데이터의 세부 속성을 분류하지 않도록 함으로써 학습시간을 크게 줄일 수 있었다는 것입니다. 또한 모델이 공격 속성을 정확히 탐지할 수 있는지 평가하기 위한 매트릭으로 일부만 정탐이라는 개념을 새롭게 정의하였습니다. 이것은 계층화된 하위의 분류를 모두 정확하게 식별한 탐지만 정탐으로 평가하도록 기준을 강화하였음을 의미합니다.

    본 논문에서는 다양한 실험 데이터 환경에서 실시간 환경에 적용할 수 있도록 높은 탐지율과 빠른 응답시간을 갖도록 하는 세 가지 향상된 성능의 침입탐지 모델을 제시하였습니다. 그리고 이 모델은 네트워크, 멀웨어, 차량 네트워크 분야의 데이터를 훈련하고 평가함으로써 탐지 성능의 향상을 검증하였습니다. 이러한 연구는 차량의 통신과 같이 실시간 탐지가 중요한 융합환경의 침입탐지 모델의 실용적인 대안이 될 수 있습니다.
    번역하기

    융합 기술의 발전으로 이기종 산업군의 시스템과 데이터는 네트워크로 상호 연결되고 있습니다. 커넥티비티 속성을 갖는 자율주행 자동차는 지능형 교통체계와 연계하여 도로 위의 혼잡지...

    융합 기술의 발전으로 이기종 산업군의 시스템과 데이터는 네트워크로 상호 연결되고 있습니다. 커넥티비티 속성을 갖는 자율주행 자동차는 지능형 교통체계와 연계하여 도로 위의 혼잡지역을 실시간으로 회피하고, 사람과 가장 가까운 거리에서 24시간 이동통신망에 연결되는 스마트 기기와 웨어러블 기기는 심박 수, 활동량, 신체리듬 등의 건강 데이터를 수집하고 분석하여 다양한 형태의 헬스케어 서비스를 제공합니다. 이러한 기기간 융합은 사용자에게 편리함을 주지만, 동시에 해커에게는 공격의 접점을 제공할 수 있습니다. 날로 지능화되고 있는 사이버 공격을 사전에 원천적으로 차단하는 것은 사실상 불가능하기 때문에 오히려 빠른 탐지로 신속하게 조치방안을 찾는 것이 더 중요해지고 있습니다. 실시간 속성이 중요한 차량의 침입탐지에서는 제한된 시간 안에 오탐이나 탐지누락을 최소화하여 유입되는 위협을 식별하고 차단하는 것이 중요합니다. 이를 위해 본 논문은 차량의 통신에서 이상 행위를 탐지하는 모델의 성능 향상을 위해 머신러닝을 사용한 세 가지 주요 분석 모델을 제안합니다.

    •우리는 데이터셋의 전처리 단계에서 모델의 입력으로 사용하기 위한 필수 특징을 추출하기 위해 correlation-entropy feature selecion (CEFS) 방법을 제안합니다. 모델은 데이터 연관성과 가중치를 기반으로 추출한 특징을 활용하여 데이터 샘플을 학습하고, 그것이 정상인지 공격인지 판별합니다. 또한 하이퍼파라미터를 조정하는 반복된 학습을 통해 탐지 성능이 향상된 모델을 도출할 수 있습니다.
    •본 논문은 데이터의 대부분이 정상 샘플을 차지하는 불균형한 양상을 띄는 트래픽에서 극소수의 공격 데이터를 놓치지 않고 탐지하기 위해 오버샘플링과 언더샘플링을 조합한 샘플링 알고리즘 combined oversampling and undersampling based on slow-start (COUSS)를 제안합니다. 그 결과 적은 수를 갖는 공격 데이터가 탐지에서 누락되지 않고, 기존보다 높은 정확도를 도출하는 것을 확인할 수 있었습니다.
    •이 논문은 공격 샘플의 상세한 속성을 분류하기 위해 다중 레이블과 다중 클래스의 구조를 사용하는 계층적 분류 알고리즘을 제안합니다. 기존의 탐지 모델이 일반적으로 대상 데이터에 공격의 발생여부 즉, 공격인지 정상인지만 분류한다면, 본 논문의 방법은 공격의 발생여부 뿐만 아니라, 공격의 종류와 해당 공격이 발생한 차량의 종류, 추가적인 공격 속성 등 세부 공격 레이블을 빠르게 식별할 수 있음을 의미합니다. 본 모델의 기여는 대부분의 정상 데이터의 세부 속성을 분류하지 않도록 함으로써 학습시간을 크게 줄일 수 있었다는 것입니다. 또한 모델이 공격 속성을 정확히 탐지할 수 있는지 평가하기 위한 매트릭으로 일부만 정탐이라는 개념을 새롭게 정의하였습니다. 이것은 계층화된 하위의 분류를 모두 정확하게 식별한 탐지만 정탐으로 평가하도록 기준을 강화하였음을 의미합니다.

    본 논문에서는 다양한 실험 데이터 환경에서 실시간 환경에 적용할 수 있도록 높은 탐지율과 빠른 응답시간을 갖도록 하는 세 가지 향상된 성능의 침입탐지 모델을 제시하였습니다. 그리고 이 모델은 네트워크, 멀웨어, 차량 네트워크 분야의 데이터를 훈련하고 평가함으로써 탐지 성능의 향상을 검증하였습니다. 이러한 연구는 차량의 통신과 같이 실시간 탐지가 중요한 융합환경의 침입탐지 모델의 실용적인 대안이 될 수 있습니다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Data and system interfaces from different types of industries can be interconnected using advanced convergence technologies. Modern vehicles connected to intelligent transport systems can avoid congested areas and select alternative routes in real-time on the road. Smart and wearable devices connected to a 24-hour mobile communication network closest to humans can collect and analyze health data such as heart rate, activity volume, and body rhythm to provide various healthcare services. This convergence between devices can provide convenience to users while providing attack surfaces to hackers. Likewise, in vehicular communications, rapid detection to find quick countermeasures become more important because it is almost impossible to block proactively all cyber-attacks that also become intelligent. An important consideration in real-time anomaly detection for vehicles is identifying and preventing incoming threats by effectively minimizing false alarms and missing detections within a limited time. This thesis proposes three major analysis methods using machine learning models to improve model performance for anomaly behaviors detection in vehicular communications.

    •As the first method, this thesis proposes a feature selection algorithm that enables the model to extract essential features by combining correlation and entropy methods in the data preprocessing stage. The model learns data samples based on the correlation and weights and classifies them as benign or malicious. Further, it can identify a model that derives improved detection performance from iterative learning to adjust the hyperparameters.
    •Further, this thesis also proposes a combined oversampling and undersampling algorithm to cover a small amount of the attack data in most benign data traffic. As a result, a small number of attacks are not missed even with the extremely imbalanced dataset. In addition, the model results in higher detection accuracy as an output than before.
    •Lastly, this thesis presents a hierarchical classification algorithm using multi-label and multiclass structures to deal with detailed attack attributes for attack samples. Conventional methods classify the target data as only a binary type to an attack or not. In contrast, our method implies that the detailed attack labels of the target data can be identified whether it is an attack or not, and the specific attributes such as attack types, vehicle types, and additional information for attack samples. Focusing on our method is the newly defined partial true positive in the evaluation metric to accurately classify which attributes specifically included in the attacks detected by the model. This means that the criteria become more strict such that only detections in the model that identify all sub-attributes accurately are evaluated as true positives.

    This paper presents three models of improved performance to ensure higher detection rates and reduced response times in various experimental datasets. In addition, these models verified improved results from simulations using three datasets in the fields of intrusions, malware, and message injections on vehicular communications. We hope that these studies can be a practical alternative model that provides meaningful anomaly detection in the vehicle and various converged real-time environments with overflowing traffic and complex intrusion scenarios.
    번역하기

    Data and system interfaces from different types of industries can be interconnected using advanced convergence technologies. Modern vehicles connected to intelligent transport systems can avoid congested areas and select alternative routes in real-tim...

    Data and system interfaces from different types of industries can be interconnected using advanced convergence technologies. Modern vehicles connected to intelligent transport systems can avoid congested areas and select alternative routes in real-time on the road. Smart and wearable devices connected to a 24-hour mobile communication network closest to humans can collect and analyze health data such as heart rate, activity volume, and body rhythm to provide various healthcare services. This convergence between devices can provide convenience to users while providing attack surfaces to hackers. Likewise, in vehicular communications, rapid detection to find quick countermeasures become more important because it is almost impossible to block proactively all cyber-attacks that also become intelligent. An important consideration in real-time anomaly detection for vehicles is identifying and preventing incoming threats by effectively minimizing false alarms and missing detections within a limited time. This thesis proposes three major analysis methods using machine learning models to improve model performance for anomaly behaviors detection in vehicular communications.

    •As the first method, this thesis proposes a feature selection algorithm that enables the model to extract essential features by combining correlation and entropy methods in the data preprocessing stage. The model learns data samples based on the correlation and weights and classifies them as benign or malicious. Further, it can identify a model that derives improved detection performance from iterative learning to adjust the hyperparameters.
    •Further, this thesis also proposes a combined oversampling and undersampling algorithm to cover a small amount of the attack data in most benign data traffic. As a result, a small number of attacks are not missed even with the extremely imbalanced dataset. In addition, the model results in higher detection accuracy as an output than before.
    •Lastly, this thesis presents a hierarchical classification algorithm using multi-label and multiclass structures to deal with detailed attack attributes for attack samples. Conventional methods classify the target data as only a binary type to an attack or not. In contrast, our method implies that the detailed attack labels of the target data can be identified whether it is an attack or not, and the specific attributes such as attack types, vehicle types, and additional information for attack samples. Focusing on our method is the newly defined partial true positive in the evaluation metric to accurately classify which attributes specifically included in the attacks detected by the model. This means that the criteria become more strict such that only detections in the model that identify all sub-attributes accurately are evaluated as true positives.

    This paper presents three models of improved performance to ensure higher detection rates and reduced response times in various experimental datasets. In addition, these models verified improved results from simulations using three datasets in the fields of intrusions, malware, and message injections on vehicular communications. We hope that these studies can be a practical alternative model that provides meaningful anomaly detection in the vehicle and various converged real-time environments with overflowing traffic and complex intrusion scenarios.

    더보기

    목차 (Table of Contents)

    • Abstract
    • Contents
    • List of Figures
    • List of Tables
    • 1 Introduction 1
    • Abstract
    • Contents
    • List of Figures
    • List of Tables
    • 1 Introduction 1
    • 1.1 Contributions 6
    • 1.2 Outline of the Thesis 10
    • 2 Background and Related Work 12
    • 2.1 Target Environments 12
    • 2.1.1 Network Attacks 12
    • 2.1.2 Mobile Malware 14
    • 2.1.3 In-Vehicle Attacks 17
    • 2.2 Data Analytics Process 21
    • 2.2.1 Preprocessing 22
    • 2.3 Related Work 27
    • 3 Network Attack Detection using Deep Learning Model 34
    • 3.1 Overview 34
    • 3.2 Network Data Analysis 35
    • 3.3 Proposed Scheme: Optimization Approach 36
    • 3.4 Simulation Results 37
    • 3.5 Summary 39
    • 4 Malware Detection using Machine Learning Model 40
    • 4.1 Overview 40
    • 4.2 Malware Data Analysis 42
    • 4.3 Proposed Scheme: Correlation-Entropy Feature Selection 43
    • 4.4 Simulation Results 48
    • 4.5 Summary 56
    • 5 Combined Oversampling and Undersampling for Imbalanced Data 57
    • 5.1 Overview 57
    • 5.2 Imbalanced Data Analysis 60
    • 5.2.1 Dataset Overview 61
    • 5.2.2 Imbalanced Ratio 64
    • 5.3 Proposed Scheme: Oversampling and Undersampling Approach 68
    • 5.4 Simulation Results 74
    • 5.5 Summary 82
    • 6 In-Vehicle Message Injection Detection using Machine Learning 83
    • 6.1 Overview 83
    • 6.2 In-Vehicle Network Intrusion Data Analysis 86
    • 6.2.1 CAN Message Frame and Topology 86
    • 6.2.2 CAN Intrusion Dataset 88
    • 6.3 Proposed Scheme: Multi-Labeled Hierarchical Classification Approach 89
    • 6.3.1 Preprocessing 89
    • 6.3.2 Model 94
    • 6.3.3 Algorithm 96
    • 6.3.4 Confusion Matrix 97
    • 6.3.5 Hypothesis Space 101
    • 6.4 Simulation 106
    • 6.4.1 Simulation Environments 106
    • 6.4.2 Simulation Results 107
    • 6.5 Summary 111
    • 7 Conclusion 113
    • Bibliography 117
    • 국문초록
    • Acknowledgement
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼