RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    크래시 마커 기반 로그 슬라이싱과 RAG 기반 코드 검색을 이용한 운영 로그 기반 분산 시스템 버그 위치 추론 기법 = Bug Localization in Distributed Systems via Crash-Marker Log Slicing and RAG-Based Code Retrieval

    한글로보기

    https://www.riss.kr/link?id=T17393314

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Diagnosing failures in large-scale distributed systems often relies solely on operational logs, as failures are difficult to reproduce and stack traces are frequently missing or unreliable. Existing bug localization techniques assume the availability of structured bug reports, making them unsuitable for real-world production settings where only raw logs remain. This thesis presents a log-based bug localization framework that integrates crash-marker log slicing, semantic code retrieval using vector embeddings, and structured reasoning with large language models (LLMs). We construct a dataset of 274 log–bug pairs from 13 open-source distributed systems and evaluate the proposed method using standard retrieval metrics. The framework achieves a file-level Recall@10 of 51.5%, demonstrating that meaningful bug localization is possible even without bug reports or stack traces. Direct LLM prompting with raw logs fails due to token limits and irrelevant context, confirming the necessity of combining log slicing and retrieval. These results provide empirical evidence that operational logs alone can support automated bug localization and lay the groundwork for log-centric debugging and reliability engineering in modern distributed systems.
    번역하기

    Diagnosing failures in large-scale distributed systems often relies solely on operational logs, as failures are difficult to reproduce and stack traces are frequently missing or unreliable. Existing bug localization techniques assume the availability ...

    Diagnosing failures in large-scale distributed systems often relies solely on operational logs, as failures are difficult to reproduce and stack traces are frequently missing or unreliable. Existing bug localization techniques assume the availability of structured bug reports, making them unsuitable for real-world production settings where only raw logs remain. This thesis presents a log-based bug localization framework that integrates crash-marker log slicing, semantic code retrieval using vector embeddings, and structured reasoning with large language models (LLMs). We construct a dataset of 274 log–bug pairs from 13 open-source distributed systems and evaluate the proposed method using standard retrieval metrics. The framework achieves a file-level Recall@10 of 51.5%, demonstrating that meaningful bug localization is possible even without bug reports or stack traces. Direct LLM prompting with raw logs fails due to token limits and irrelevant context, confirming the necessity of combining log slicing and retrieval. These results provide empirical evidence that operational logs alone can support automated bug localization and lay the groundwork for log-centric debugging and reliability engineering in modern distributed systems.

    더보기

    목차 (Table of Contents)

    • 1. 개 요 1
    • 2. 논문 구성 7
    • 3. 배경 지식 9
    • 3.1. 분산시스템과 운영 로그 9
    • 3.1.1 분산 시스템 9
    • 1. 개 요 1
    • 2. 논문 구성 7
    • 3. 배경 지식 9
    • 3.1. 분산시스템과 운영 로그 9
    • 3.1.1 분산 시스템 9
    • 3.1.2 운영 로그 11
    • 3.1.3 스택 트레이스 12
    • 3.2 로그 기반 장애 진단 15
    • 3.3 버그 로컬라이제이션 18
    • 3.3.1 정보검색 기반 기법 19
    • 3.3.2 스펙트럼 기반 기법 21
    • 3.4 연구 관련 핵심 기술 22
    • 3.4.1 대형 언어모델 22
    • 3.4.2 검색 기반 생성(RAG) 23
    • 3.4.3 벡터 저장소 24
    • 4. 데이터셋 구축 및 문제 정의 27
    • 4.1. 데이터셋 구축 27
    • 4.1.1 대상 프로젝트 선정 27
    • 4.1.2 운영 로그 수집 28
    • 4.1.3 정답 데이터 구성 29
    • 4.1.4 데이터셋 통계 35
    • 4.2. 문제 정의 36
    • 4.3. 연구 질문 39
    • 5. 운영 로그 기반 버그 위치 추론 기법 43
    • 5.1. 전체 구조 43
    • 5.2. 크래시 마커 기반 로그 슬라이싱 44
    • 5.3. 코드 임베딩 및 소스 파일 검색 51
    • 5.4. LLM 기반 버그 위치 추론 52
    • 5.5. 결과 순위화 55
    • 6. 실험 설계 및 결과 57
    • 6.1. 실험 환경 57
    • 6.2. 평가 지표 57
    • 6.3. 실험 결과 59
    • 6.4. 한계점 65
    • 7. 결론 68
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼