RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    데이터센터 환경을 위한 AI 기반 네트워크 인프라 통합 관제 시스템에 관한 연구 = A Study on Integrated Monitoring System of AI?Based Network Infrastructure for Data Center Environment

    한글로보기

    https://www.riss.kr/link?id=T17405601

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    디지털 전환과 마이크로서비스·컨테이너의 확산에 힘입어 데이터센터 네트워크는 동–서(East–West) 트래픽이 지배하는 고도로 분산되고 동적인 환경으로 진화하였다. 정적 임계치와 사일로화된 지표에 의존하는 전통적 모니터링은 경고 폭주, 미세 이상 미탐, 근본 원인 식별 지연과 같은 한계가 있다.
    본 논문은 스트리밍 텔레메트리 기반에서 관측 가능성 3요소(로그·메트릭·트레이스)를 통합하고, AIOps 기법을 종단 간 적용한 AI 기반 통합 관제 시스템을 설계, 구현 및 검증한다.
    데이터 수집·처리 계층에서는 gNMI, OpenTelemetry, Fluentd를 통해 이기종 텔레메트리를 정규화하고 시간 정렬하여 Kafka 파이프라인으로 전달한 뒤, InfluxDB와 Elasticsearch에 저장한다. 분석 계층은 시간적 문맥을 학습하는 LSTM–오토인코더로 다변량 시계열의 이상을 탐지하고, Prophet으로 용량과 성능을 예측한다. 더불어 시간·토폴로지 상관을 결합한 근본 원인 분석(RCA) 워크플로가 이벤트 군집을 추론하여 운영자의 의사결정을 지원한다. 시각화 계층은 단일 창(Single Pane) 대시보드, 토폴로지 오버레이, 상관 분석 뷰를 통해 탐지–예측–분석을 일관되게 제공한다.
    평가는 공개 트래픽 데이터셋(UNSW–NB15, CIC–IDS2017)에서 추출한 플로우 통계와, GNS3 기반 Clos 토폴로지에서 주입한 네 가지 장애(인캐스트 혼잡, 포트 플래핑, 라우팅 루프, 마이크로버스트)를 결합한 하이브리드 데이터셋(4주, 10초 해상도, 12개 지표)을 사용하였다. 제안 모델은 정적 임계치(3–sigma)와 Isolation Forest 대비 정밀도 0.91, 재현율 0.94, F1–Score 0.92로 우수한 이상 탐지 성능을 보였다. 예측 측면에서 Prophet은 ARIMA 대비 MAE 89.7Mbps, RMSE 125.3Mbps로 오차를 크게 감소시켰다. RCA는 평균 95.0% 정확도로 근본 원인을 식별했으며, 근본 원인 특정까지의 MTTR을 26.8분에서 4.4분으로 약 86% 단축하였다. 복합 장애 및 마이크로버스트와 같은 잠재적 성능 저하 시나리오에서도 일관된 효과가 확인되었다.

    본 연구의 기여는 다음과 같다.
    ⑴ 로그–메트릭–트레이스를 아우르는 통합 관측 파이프라인 설계
    ⑵ LSTM 오토인코더와 Prophet을 결합한 이상 탐지–예측 엔진 제안
    ⑶ 시간·토폴로지 상관 기반 RCA 워크플로와 단일 창 시각화
    ⑷ 실환경에 근접한 하이브리드 데이터셋을 통한 정량 검증

    제안 시스템은 반응적 모니터링을 예측적·자동화된 운영으로 전환하는 실용적 경로를 제시하였고, 향후 모델 경량화, 엣지 추론, 자율 복구(Self–Healing)로의 확장 가능성을 논의한다.
    번역하기

    디지털 전환과 마이크로서비스·컨테이너의 확산에 힘입어 데이터센터 네트워크는 동–서(East–West) 트래픽이 지배하는 고도로 분산되고 동적인 환경으로 진화하였다. 정적 임계치와 사일...

    디지털 전환과 마이크로서비스·컨테이너의 확산에 힘입어 데이터센터 네트워크는 동–서(East–West) 트래픽이 지배하는 고도로 분산되고 동적인 환경으로 진화하였다. 정적 임계치와 사일로화된 지표에 의존하는 전통적 모니터링은 경고 폭주, 미세 이상 미탐, 근본 원인 식별 지연과 같은 한계가 있다.
    본 논문은 스트리밍 텔레메트리 기반에서 관측 가능성 3요소(로그·메트릭·트레이스)를 통합하고, AIOps 기법을 종단 간 적용한 AI 기반 통합 관제 시스템을 설계, 구현 및 검증한다.
    데이터 수집·처리 계층에서는 gNMI, OpenTelemetry, Fluentd를 통해 이기종 텔레메트리를 정규화하고 시간 정렬하여 Kafka 파이프라인으로 전달한 뒤, InfluxDB와 Elasticsearch에 저장한다. 분석 계층은 시간적 문맥을 학습하는 LSTM–오토인코더로 다변량 시계열의 이상을 탐지하고, Prophet으로 용량과 성능을 예측한다. 더불어 시간·토폴로지 상관을 결합한 근본 원인 분석(RCA) 워크플로가 이벤트 군집을 추론하여 운영자의 의사결정을 지원한다. 시각화 계층은 단일 창(Single Pane) 대시보드, 토폴로지 오버레이, 상관 분석 뷰를 통해 탐지–예측–분석을 일관되게 제공한다.
    평가는 공개 트래픽 데이터셋(UNSW–NB15, CIC–IDS2017)에서 추출한 플로우 통계와, GNS3 기반 Clos 토폴로지에서 주입한 네 가지 장애(인캐스트 혼잡, 포트 플래핑, 라우팅 루프, 마이크로버스트)를 결합한 하이브리드 데이터셋(4주, 10초 해상도, 12개 지표)을 사용하였다. 제안 모델은 정적 임계치(3–sigma)와 Isolation Forest 대비 정밀도 0.91, 재현율 0.94, F1–Score 0.92로 우수한 이상 탐지 성능을 보였다. 예측 측면에서 Prophet은 ARIMA 대비 MAE 89.7Mbps, RMSE 125.3Mbps로 오차를 크게 감소시켰다. RCA는 평균 95.0% 정확도로 근본 원인을 식별했으며, 근본 원인 특정까지의 MTTR을 26.8분에서 4.4분으로 약 86% 단축하였다. 복합 장애 및 마이크로버스트와 같은 잠재적 성능 저하 시나리오에서도 일관된 효과가 확인되었다.

    본 연구의 기여는 다음과 같다.
    ⑴ 로그–메트릭–트레이스를 아우르는 통합 관측 파이프라인 설계
    ⑵ LSTM 오토인코더와 Prophet을 결합한 이상 탐지–예측 엔진 제안
    ⑶ 시간·토폴로지 상관 기반 RCA 워크플로와 단일 창 시각화
    ⑷ 실환경에 근접한 하이브리드 데이터셋을 통한 정량 검증

    제안 시스템은 반응적 모니터링을 예측적·자동화된 운영으로 전환하는 실용적 경로를 제시하였고, 향후 모델 경량화, 엣지 추론, 자율 복구(Self–Healing)로의 확장 가능성을 논의한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Driven by digital transformation and the proliferation of microservices and containers, data–center networks have evolved into highly distributed and dynamic environments dominated by east–west traffic. Traditional monitoring, which relies on static thresholds and siloed indicators, suffers from alert storms, missed subtle anomalies, and delays in root–cause identification.
    This dissertation designs, implements, and evaluates an AI–based integrated monitoring system that unifies the three pillars of observability (logs, metrics, and traces) over streaming telemetry, and applies AIOps techniques in an end–to–end manner.
    In the data collection and processing layer, heterogeneous telemetry is normalized and time–aligned via gNMI, OpenTelemetry, and Fluentd, streamed through a Kafka pipeline, and stored in InfluxDB and Elasticsearch. The analytics layer employs an LSTM–autoencoder to detect anomalies in multivariate time series by learning temporal context, and uses Prophet to forecast capacity and performance. In addition, a root–cause analysis (RCA) workflow that combines temporal and topological correlations infers event clusters and supports operator decision–making. The visualization layer consistently delivers detection, prediction, and analysis results via a single–pane dashboard, topology overlays, and correlation views.
    Evaluation is conducted on a hybrid dataset that combines flow statistics extracted from public traffic datasets (UNSW–NB15, CIC–IDS2017) with four injected faults (incast congestion, port flapping, routing loop, and microburst) on a GNS3–based Clos topology (4 weeks, 10–second resolution, 12 metrics). The proposed anomaly detection model outperforms static 3–sigma thresholding and Isolation Forest, achieving a precision of 0.91, recall of 0.94, and F1–score of 0.92. For forecasting, Prophet reduces error substantially compared with ARIMA, with MAE and RMSE improved by 89.7 Mbps and 125.3 Mbps, respectively. The RCA achieves an average accuracy of 95.0% and shortens MTTR to pinpoint the root cause from 26.8 minutes to 4.4 minutes (–86.0%). Consistent effectiveness is observed even under composite failures and latent performance–degradation scenarios such as microbursts.

    The contributions of this dissertation are as follows:
    ⑴ design of a unified observability pipeline that spans logs, metrics, and traces;
    ⑵ proposal of an integrated detection–forecast engine combining an LSTM autoencoder and Prophet;
    ⑶ development of a temporal and topology–aware RCA workflow with single–pane visualization;
    ⑷ quantitative validation using a hybrid dataset closely reflecting real data–center environments.

    The proposed system presents a practical path for transforming reactive monitoring into predictive and automated operations, and the dissertation discusses future extensions toward model lightweighting, edge inference, and self–healing capabilities.
    번역하기

    Driven by digital transformation and the proliferation of microservices and containers, data–center networks have evolved into highly distributed and dynamic environments dominated by east–west traffic. Traditional monitoring, which relies on stat...

    Driven by digital transformation and the proliferation of microservices and containers, data–center networks have evolved into highly distributed and dynamic environments dominated by east–west traffic. Traditional monitoring, which relies on static thresholds and siloed indicators, suffers from alert storms, missed subtle anomalies, and delays in root–cause identification.
    This dissertation designs, implements, and evaluates an AI–based integrated monitoring system that unifies the three pillars of observability (logs, metrics, and traces) over streaming telemetry, and applies AIOps techniques in an end–to–end manner.
    In the data collection and processing layer, heterogeneous telemetry is normalized and time–aligned via gNMI, OpenTelemetry, and Fluentd, streamed through a Kafka pipeline, and stored in InfluxDB and Elasticsearch. The analytics layer employs an LSTM–autoencoder to detect anomalies in multivariate time series by learning temporal context, and uses Prophet to forecast capacity and performance. In addition, a root–cause analysis (RCA) workflow that combines temporal and topological correlations infers event clusters and supports operator decision–making. The visualization layer consistently delivers detection, prediction, and analysis results via a single–pane dashboard, topology overlays, and correlation views.
    Evaluation is conducted on a hybrid dataset that combines flow statistics extracted from public traffic datasets (UNSW–NB15, CIC–IDS2017) with four injected faults (incast congestion, port flapping, routing loop, and microburst) on a GNS3–based Clos topology (4 weeks, 10–second resolution, 12 metrics). The proposed anomaly detection model outperforms static 3–sigma thresholding and Isolation Forest, achieving a precision of 0.91, recall of 0.94, and F1–score of 0.92. For forecasting, Prophet reduces error substantially compared with ARIMA, with MAE and RMSE improved by 89.7 Mbps and 125.3 Mbps, respectively. The RCA achieves an average accuracy of 95.0% and shortens MTTR to pinpoint the root cause from 26.8 minutes to 4.4 minutes (–86.0%). Consistent effectiveness is observed even under composite failures and latent performance–degradation scenarios such as microbursts.

    The contributions of this dissertation are as follows:
    ⑴ design of a unified observability pipeline that spans logs, metrics, and traces;
    ⑵ proposal of an integrated detection–forecast engine combining an LSTM autoencoder and Prophet;
    ⑶ development of a temporal and topology–aware RCA workflow with single–pane visualization;
    ⑷ quantitative validation using a hybrid dataset closely reflecting real data–center environments.

    The proposed system presents a practical path for transforming reactive monitoring into predictive and automated operations, and the dissertation discusses future extensions toward model lightweighting, edge inference, and self–healing capabilities.

    더보기

    목차 (Table of Contents)

    • 국문초록 ⅰ
    • 목 차 ⅲ
    • 그림목차 ⅶ
    • 표 목 차 ⅷ
    • 약 어 표 ⅸ
    • 국문초록 ⅰ
    • 목 차 ⅲ
    • 그림목차 ⅶ
    • 표 목 차 ⅷ
    • 약 어 표 ⅸ
    • Ⅰ. 서 론 1
    • 1.1 연구의 배경 및 필요성 1
    • 1.1.1 데이터센터 환경의 변화 1
    • 1.1.2 전통적인 모니터링 방식의 한계 2
    • 1.1.3 관측 가능성(Observability)과 AIOps의 대두 3
    • 1.2 연구의 목적 및 범위 4
    • 1.3 연구의 방법론 및 절차 5
    • 1.4 논문의 구성 6
    • Ⅱ. 이론적 배경 및 관련 연구 8
    • 2.1 데이터센터 네트워크 아키텍처 8
    • 2.1.1 스파인–리프(Spine–Leaf) 구조 8
    • 2.1.2 가상화 및 오버레이 네트워크(Overlay Network) 9
    • 2.2 네트워크 관제 패러다임의 진화 10
    • 2.2.1 모니터링(Monitoring)과 관측 가능성(Observability) 10
    • 2.2.2 관측 가능성의 세 가지 핵심 요소 14
    • 2.2.3 스트리밍 텔레메트리(Streaming Telemetry) 기술 18
    • 2.3 AIOps(AI for IT Operations) 21
    • 2.3.1 AIOps의 개념 및 핵심 기능 21
    • 2.3.2 AIOps 플랫폼 구성 요소 22
    • 2.4 AI 기반 시계열 데이터 분석 기술 24
    • 2.4.1 딥러닝 기반 이상 탐지 24
    • 2.4.2 시계열 예측 모델 27
    • 2.5 선행 연구 분석 및 본 연구의 차별성 28
    • Ⅲ. AI 기반 통합 관제 시스템 설계 31
    • 3.1 시스템 전체 아키텍처 31
    • 3.2 데이터 수집 및 처리 계층 설계 33
    • 3.2.1 텔레메트리 데이터 수집기 설계 34
    • 3.2.2 데이터 스트리밍 파이프라인 구축 34
    • 3.2.3 데이터 정규화 및 특징 공학 35
    • 3.3 AI 분석 엔진 설계 35
    • 3.3.1 LSTM–오토인코더 기반 이상 탐지 모델 설계 36
    • 3.3.2 Prophet 기반 장애 예측 모델 설계 37
    • 3.3.3 근본 원인 분석(RCA) 알고리즘 설계 38
    • 3.4 관제 및 시각화 대시보드 설계 40
    • 3.4.1 통합 대시보드 41
    • 3.4.2 토폴로지 기반 시각화 42
    • 3.4.3 상관관계 분석 뷰(Correlation Analysis View) 43
    • Ⅳ. 시스템 구현 및 실험 환경 45
    • 4.1 시스템 개발 환경 45
    • 4.2 실험 데이터셋 47
    • 4.2.1 데이터 소스 및 수집 47
    • 4.2.2 데이터셋 특성 및 전처리 49
    • 4.2.3 하이브리드 데이터셋 구성 과정 51
    • 4.3 성능 평가 지표 52
    • 4.3.1 이상 탐지 모델 평가 지표 52
    • 4.3.2 예측 모델 평가 지표 54
    • 4.3.3 근본 원인 분석 평가 지표 55
    • 4.4 비교 대상 모델 선정 57
    • 4.4.1 이상 탐지 기능 비교 모델 57
    • 4.4.2 시계열 예측 기능 비교 모델 59
    • 4.4.3 근본 원인 분석(RCA) 비교 대상 기법 선정 60
    • Ⅴ. 실험 결과 및 분석 62
    • 5.1 이상 탐지 모델 성능 평가 결과 62
    • 5.1.1 주요 파라미터 튜닝 결과 62
    • 5.1.2 비교 모델과의 성능 비교 분석 64
    • 5.1.3 오탐지 및 미탐지 사례 분석 65
    • 5.2 장애 예측 모델 성능 평가 결과 66
    • 5.2.1 주요 지표 예측 정확도 분석 66
    • 5.2.2 예측 기반 사전 경고 시스템의 유효성 검증 67
    • 5.3 근본 원인 분석 시나리오 기반 성능 평가 70
    • 5.3.1 시나리오 1 : 단일 장애 발생 시 RCA 성능 71
    • 5.3.2 시나리오 2 : 복합 장애 발생 시 RCA 성능 72
    • 5.3.3 MTTR(평균 해결 시간) 개선 효과 분석 73
    • 5.4 RCA 오판 사례 분석 75
    • 5.4.1 오판 사례의 유형 및 원인 분석 75
    • 5.4.2 RCA 개선 방향 76
    • 5.5 시스템 통합 성능 및 효과 고찰 79
    • Ⅵ. 결론 및 향후 연구 82
    • 6.1 연구 결과 요약 및 의의 82
    • 6.2 연구의 한계점 83
    • 6.3 향후 연구 방향 84
    • 6.3.1 모델 경량화 및 실시간 추론 성능 개선 84
    • 6.3.2 자동화된 복구(Self–Healing) 시스템으로의 확장 86
    • 참고문헌(References) 88
    • 영문초록 96
    • 감사의 글(Acknowledgement) 99
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    연관 공개강의(KOCW)

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼