RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    (A) Study on structured and explainable communication data analysis for digital forensics

    한글로보기

    https://www.riss.kr/link?id=T17387470

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    디지털 포렌식에서 커뮤니케이션 데이터(인스턴트 메신저 대화, SMS 기록, 이메일)는 행위자의 의도, 관계에서의 상호작용, 사건의 타임라인 등을 드러내는 주요 디지털 증거이다. 그러나 커뮤니케이션 데이터의 대량성, 비정형성, 이질성과 같은 특성으로 기존 분석 방법인 키워드 검색이나 규칙 기반 필터링 방식으로 복잡한 의미의 연결성이나 범죄 패턴을 포착하는 데 한계를 가진다. 최근 제시되고 있는 AI 기반 포렌식 분석 방안도 대화의 맥락, 관계 구조를 일관되게 복원하기 어려우며, 검증 불가능한 내용을 생성하고(hallucination), 추론의 경로를 명확히 드러내지 못하며, 디지털 증거가 갖추어야할 설명 가능성과 증거능력을 보장하지 못한다.
    이러한 한계를 해결하기 위해, 본 연구는 커뮤니케이션 데이터를 대상으로 한 그래프 기반 검색-증강 생성 프레임워크인 DF-Graph(Digital Forensics-GraphRAG)를 제안한다. DF-Graph는 메시지 로그를 기준으로 노드–엣지(node-edge) 기반 지식 그래프 형태로 구조화하고, 의미적·구조적 단서를 활용하여 질의와 관련된 서브그래프를 검색한다. 이후 포렌식에 특화된 프롬프트를 통해 답변을 생성하며, 규칙 기반 추론 경로(rule-based reasoning trace)와 디지털 증거로 인정될만한 수준의 인용을 통해 높은 신뢰성을 제공한다.
    또한 메시지 로그의 대부분을 차지하는 텍스트 데이터 분석 프레임워크를 구현하고 나아가 다양한 모달리티(이미지, 음성)도 데이터 소스로 추가하여 새로운 분석 프레임워크의 활용 가능성을 확인 한다.
    본 연구에서는 실제 데이터, 공개 데이터, 그리고 『죄와 벌(Crime and Punishment)』에서 서사 구조를 변형하여 생성한 합성 데이터 등 다양한 환경에서 DF-Graph를 평가하였다.
    비교 대상으로는 (1) GPT 기반 직접 생성 방식(GPT-only), (2) BERT 기반 의도 분류와 GPT를 결합한 하이브리드 방식(BERT+GPT), (3) dense retrieval 기반 단순 RAG 방식(Naive RAG), (4) 제안하는 그래프 기반 검색 방식(DF-Graph)이 포함된다.
    실험 결과, DF-Graph는 정확도(Accuracy) 57.23%, 의미 유사도(BERTScore F1) 0.8597, 문맥 충실도(Faithfulness) 5.561 등 모든 지표에서 일관되게 가장 높은 성능을 달성하였다. 또한 8명의 디지털 포렌식 전문가를 대상으로 한 사용자 연구에서도 DF-Graph는 기존 방식보다 더 설명 가능하고, 정확하며, 법적으로 방어 가능한 결과를 제공한다는 평가를 받았다.
    또한, 멀티모달 통합 분석 프레임워크(MM-DF-Graph)를 제안하여 텍스트 중심 분석만으로는 확인되지 않던 증거가 멀티모달 정보(이미지·음성 파일)의 통합을 통해 효과적으로 보강될 수 있음을 확인하였다. 기존 텍스트 중심 그래프구조를 변경하지 않고 멀티모달 데이터를 하나의 메시지 단위로 정규화 하여 맥락 보존을 강화하고 분석 정확도와 설명 가능성을 향상시키는 새로운 통합 분석 프레임워크를 최초로 제안 한다.
    이로써 본 연구에서 제시한 DF-Graph와 MM-DF-Graph가 AI 기반 디지털 포렌식 도구로 발전할 수 있는 실용적 가능성과 신뢰성을 갖추고 있음을 보여주는 근거가 된다. 본 논문은 향후 디지털 포렌식 분석 프레임워크가 텍스트뿐 아니라 다양한 형태의 디지털 증거를 통합적으로 다루는 방향으로 확장될 수 있는 방향을 제시하며, AI기반 디지털 포렌식 연구의 기초를 마련한다.
    번역하기

    디지털 포렌식에서 커뮤니케이션 데이터(인스턴트 메신저 대화, SMS 기록, 이메일)는 행위자의 의도, 관계에서의 상호작용, 사건의 타임라인 등을 드러내는 주요 디지털 증거이다. 그러나 커...

    디지털 포렌식에서 커뮤니케이션 데이터(인스턴트 메신저 대화, SMS 기록, 이메일)는 행위자의 의도, 관계에서의 상호작용, 사건의 타임라인 등을 드러내는 주요 디지털 증거이다. 그러나 커뮤니케이션 데이터의 대량성, 비정형성, 이질성과 같은 특성으로 기존 분석 방법인 키워드 검색이나 규칙 기반 필터링 방식으로 복잡한 의미의 연결성이나 범죄 패턴을 포착하는 데 한계를 가진다. 최근 제시되고 있는 AI 기반 포렌식 분석 방안도 대화의 맥락, 관계 구조를 일관되게 복원하기 어려우며, 검증 불가능한 내용을 생성하고(hallucination), 추론의 경로를 명확히 드러내지 못하며, 디지털 증거가 갖추어야할 설명 가능성과 증거능력을 보장하지 못한다.
    이러한 한계를 해결하기 위해, 본 연구는 커뮤니케이션 데이터를 대상으로 한 그래프 기반 검색-증강 생성 프레임워크인 DF-Graph(Digital Forensics-GraphRAG)를 제안한다. DF-Graph는 메시지 로그를 기준으로 노드–엣지(node-edge) 기반 지식 그래프 형태로 구조화하고, 의미적·구조적 단서를 활용하여 질의와 관련된 서브그래프를 검색한다. 이후 포렌식에 특화된 프롬프트를 통해 답변을 생성하며, 규칙 기반 추론 경로(rule-based reasoning trace)와 디지털 증거로 인정될만한 수준의 인용을 통해 높은 신뢰성을 제공한다.
    또한 메시지 로그의 대부분을 차지하는 텍스트 데이터 분석 프레임워크를 구현하고 나아가 다양한 모달리티(이미지, 음성)도 데이터 소스로 추가하여 새로운 분석 프레임워크의 활용 가능성을 확인 한다.
    본 연구에서는 실제 데이터, 공개 데이터, 그리고 『죄와 벌(Crime and Punishment)』에서 서사 구조를 변형하여 생성한 합성 데이터 등 다양한 환경에서 DF-Graph를 평가하였다.
    비교 대상으로는 (1) GPT 기반 직접 생성 방식(GPT-only), (2) BERT 기반 의도 분류와 GPT를 결합한 하이브리드 방식(BERT+GPT), (3) dense retrieval 기반 단순 RAG 방식(Naive RAG), (4) 제안하는 그래프 기반 검색 방식(DF-Graph)이 포함된다.
    실험 결과, DF-Graph는 정확도(Accuracy) 57.23%, 의미 유사도(BERTScore F1) 0.8597, 문맥 충실도(Faithfulness) 5.561 등 모든 지표에서 일관되게 가장 높은 성능을 달성하였다. 또한 8명의 디지털 포렌식 전문가를 대상으로 한 사용자 연구에서도 DF-Graph는 기존 방식보다 더 설명 가능하고, 정확하며, 법적으로 방어 가능한 결과를 제공한다는 평가를 받았다.
    또한, 멀티모달 통합 분석 프레임워크(MM-DF-Graph)를 제안하여 텍스트 중심 분석만으로는 확인되지 않던 증거가 멀티모달 정보(이미지·음성 파일)의 통합을 통해 효과적으로 보강될 수 있음을 확인하였다. 기존 텍스트 중심 그래프구조를 변경하지 않고 멀티모달 데이터를 하나의 메시지 단위로 정규화 하여 맥락 보존을 강화하고 분석 정확도와 설명 가능성을 향상시키는 새로운 통합 분석 프레임워크를 최초로 제안 한다.
    이로써 본 연구에서 제시한 DF-Graph와 MM-DF-Graph가 AI 기반 디지털 포렌식 도구로 발전할 수 있는 실용적 가능성과 신뢰성을 갖추고 있음을 보여주는 근거가 된다. 본 논문은 향후 디지털 포렌식 분석 프레임워크가 텍스트뿐 아니라 다양한 형태의 디지털 증거를 통합적으로 다루는 방향으로 확장될 수 있는 방향을 제시하며, AI기반 디지털 포렌식 연구의 기초를 마련한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Communication data, such as instant messenger exchanges, SMS records, and emails, plays a critical role in digital forensic investigations by revealing criminal intent, interpersonal dynamics, and the temporal structure of events. However, existing AI-based forensic tools frequently hallucinate unverifiable content, obscure their reasoning paths, and ultimately fail to meet the traceability and legal admissibility standards required in criminal investigations. To overcome these challenges, we propose DF-Graph, a graph-based retrieval-augmented generation (Graph-RAG) framework designed for forensic question answering over communication data. DF-Graph constructs structured knowledge graphs from message logs, retrieves query-relevant subgraphs based on semantic and structural cues, and generates answers guided by forensic-specific prompts. It further enhances legal transparency through rule-based reasoning traces and citation of message level evidence. We comprehensively evaluate DF-Graph across real-world, public, and synthetic datasets, including a narrative dataset adapted from Crime and Punishment. Our evaluation compares four approaches: (1) a direct generation approach using only a language model without retrieval; (2) a BERT embedding-based selective retrieval approach that identifies relevant messages before generation; (3) a conventional text-based retrieval approach; and (4) our proposed graph-based retrieval approach (DF-Graph). Empirical results show that DF-Graph consistently outperforms all baseline approaches in Accuracy (57.23 %), semantic similarity (BERTScore F1: 0.8597), and contextual faithfulness. A user study with eight forensic experts confirms that DF-Graph delivers more explainable, accurate, and legally defensible outputs, making it a practical solution for AI assisted forensic investigations.
    Furthermore, a multimodal integration framework is introduced, demonstrating that evidence undetectable through text only analysis can be effectively recovered When image and audio data are unified within the analysis pipeline. By normalizing multimodal inputs into a single message level representation without modifying the underlying text-centric graph structure, the framework enhances contextual preservation, analytical accuracy, and forensic explainability.
    Collectively, these findings demonstrate that DF-Graph and the multimodal integration framework offer practical and reliable foundations for next-generation AI-assisted digital forensic tools. This study suggests a pathway toward forensic analysis systems capable of handling diverse forms of digital evidence beyond text, establishing an initial foundation for multimodal AI-driven digital forensics research.
    번역하기

    Communication data, such as instant messenger exchanges, SMS records, and emails, plays a critical role in digital forensic investigations by revealing criminal intent, interpersonal dynamics, and the temporal structure of events. However, existing AI...

    Communication data, such as instant messenger exchanges, SMS records, and emails, plays a critical role in digital forensic investigations by revealing criminal intent, interpersonal dynamics, and the temporal structure of events. However, existing AI-based forensic tools frequently hallucinate unverifiable content, obscure their reasoning paths, and ultimately fail to meet the traceability and legal admissibility standards required in criminal investigations. To overcome these challenges, we propose DF-Graph, a graph-based retrieval-augmented generation (Graph-RAG) framework designed for forensic question answering over communication data. DF-Graph constructs structured knowledge graphs from message logs, retrieves query-relevant subgraphs based on semantic and structural cues, and generates answers guided by forensic-specific prompts. It further enhances legal transparency through rule-based reasoning traces and citation of message level evidence. We comprehensively evaluate DF-Graph across real-world, public, and synthetic datasets, including a narrative dataset adapted from Crime and Punishment. Our evaluation compares four approaches: (1) a direct generation approach using only a language model without retrieval; (2) a BERT embedding-based selective retrieval approach that identifies relevant messages before generation; (3) a conventional text-based retrieval approach; and (4) our proposed graph-based retrieval approach (DF-Graph). Empirical results show that DF-Graph consistently outperforms all baseline approaches in Accuracy (57.23 %), semantic similarity (BERTScore F1: 0.8597), and contextual faithfulness. A user study with eight forensic experts confirms that DF-Graph delivers more explainable, accurate, and legally defensible outputs, making it a practical solution for AI assisted forensic investigations.
    Furthermore, a multimodal integration framework is introduced, demonstrating that evidence undetectable through text only analysis can be effectively recovered When image and audio data are unified within the analysis pipeline. By normalizing multimodal inputs into a single message level representation without modifying the underlying text-centric graph structure, the framework enhances contextual preservation, analytical accuracy, and forensic explainability.
    Collectively, these findings demonstrate that DF-Graph and the multimodal integration framework offer practical and reliable foundations for next-generation AI-assisted digital forensic tools. This study suggests a pathway toward forensic analysis systems capable of handling diverse forms of digital evidence beyond text, establishing an initial foundation for multimodal AI-driven digital forensics research.

    더보기

    목차 (Table of Contents)

    • Ⅰ 서론 1
    • 1. 연구의 배경 및 목적 1
    • 1.1. 연구의 배경 1
    • 1.2. 연구의 목적 2
    • 2. 연구의 범위 및 방법 3
    • Ⅰ 서론 1
    • 1. 연구의 배경 및 목적 1
    • 1.1. 연구의 배경 1
    • 1.2. 연구의 목적 2
    • 2. 연구의 범위 및 방법 3
    • 2.1. 연구의 범위 3
    • 2.2. 연구 방법 4
    • Ⅱ 관련연구 6
    • 1. 디지털 포렌식의 기본원칙과 변화 6
    • 2. 디지털 포렌식 분야에서의 인공지능 기술 적용 7
    • 3. 디지털 포렌식에서의 RAG와 설명 가능성 9
    • 4. 커뮤니케이션 데이터 분석 10
    • 4.1. 커뮤니케이션 데이터의 중요성 10
    • 4.2. 단일모달 분석 : 텍스트 기반 분석 11
    • 4.3. 멀티모달 분석 12
    • Ⅲ 문제 정의 13
    • 1. 커뮤니케이션 데이터의 구조적 특성 13
    • 2. LLM, Naive RAG 적용의 한계 17
    • 3. 문제 정의 18
    • Ⅳ DF-Graph 20
    • 1. 개요 20
    • 1.1. 프레임워크 구성 요소 20
    • 1.2. 프레임워크 설계 원칙 21
    • 2. 프레임워크 아키텍처 23
    • 2.1. 전처리 단계 23
    • 2.1.1.텍스트 추출 24
    • 2.1.2. 구조 정규화 24
    • 2.1.3. 시간 순서 정렬 및 정규화 25
    • 2.1.4. 개인정보 익명화 및 민감정보 비식별화 26
    • 2.2. 그래프 구축 단계 26
    • 2.2.1. 메시지 단위 그래프 모델 정의 26
    • 2.2.2. 그래프 노드 구성 27
    • 2.2.3. 그래프 엣지 구성 28
    • 2.2.4. 그래프 클러스터링 및 커뮤니티 구조 32
    • 2.2.5. 한계 및 재현성 문제 33
    • 2.3. 서브그래프 구축 단계 33
    • 2.3.1. 서브그래프의 의미 33
    • 2.3.2. 의미 기반 후보 노드 탐색 34
    • 2.3.3. 그래프 확장 34
    • 2.3.4. 서브그래프의 특징과 역할 35
    • 2.4. 근거 기반 응답 생성 단계 35
    • 2.4.1. 증거 컨텍스트 선형화 35
    • 2.4.2. 포렌식 질의 프롬프트 구성 36
    • 2.5. 설명 가능한 추론 경로 단계 39
    • 2.6. 구현 상세 40
    • 2.6.1. 그래프 엔진 및 저장 구조 40
    • 2.6.2. 의미 임베딩 및 검색 40
    • 2.6.3. 프롬프트 생성 41
    • 2.6.4. 언어 모델 백엔드 41
    • 2.6.5. 추적 가능성 모듈 41
    • 2.6.6. 실험 환경 42
    • 3. 실험 및 평가 43
    • 3.1. 정량 평가 43
    • 3.1.1. 비교모델 43
    • 3.1.2. 데이터 소스 및 준비 과정 47
    • 3.1.3. 평가 지표 50
    • 3.1.4. 실험 결과 52
    • 3.1.5. 통계적 타당성 검증 53
    • 3.2. 사용자 연구 57
    • 3.2.1. 연구 설계 및 평가 57
    • 3.2.2. 참여자 58
    • 3.2.3. 결과 59
    • 4. 논의 및 결론 62
    • 4.1. 포렌식 추론 체계화와의 정합성 62
    • 4.2. 추적 가능성과 법적 투명성 64
    • 4.3. 실무 활용을 위한 고려사항 64
    • 4.3.1. 확장성 65
    • 4.3.2. 불확실성 처리 65
    • 4.3.3. 인프라 제약 65
    • Ⅴ MM-DF-Graph 66
    • 1. 개요 66
    • 1.1. 멀티모달 통합 분석 프레임워크의 필요성 66
    • 1.2. 프레임워크 구성 요소설계 원칙 67
    • 1.3. 프레임워크 설계 원칙 68
    • 2. 프레임워크 구성 69
    • 2.1. 멀티모달 데이터 추출 단계 69
    • 2.2. 텍스트 스크립팅 단계 70
    • 2.2.1. 이미지 데이터 처리 71
    • 2.2.2. 음성 데이터 처리 73
    • 2.3. 정규화 단계 73
    • 2.4. 그래프 생성 단계 75
    • 2.5. 질의응답 및 검증 단계 75
    • 3. 실험 및 평가 78
    • 3.1. 개요 78
    • 3.2. 실험 78
    • 3.2.1. 그래프 엔진 및 저장 구조 78
    • 3.2.2. 멀티모달 전처리 79
    • 3.2.3. 임베딩 및 검색 79
    • 3.2.4. 프롬프트 생성 79
    • 3.2.5. 언어모델 백엔드 79
    • 3.2.6. 실험 환경 79
    • 3.3. 평가 80
    • 3.3.1. 데이터셋 80
    • 3.3.2. 사례 기반 비교 81
    • 3.3.3. 포렌식 관점에서의 구조적 평가 84
    • 3.3.4. Ablation Study 85
    • 4. 논의 및 결과 87
    • 4.1. 멀티모달 통합 분석 가능성 87
    • 4.2. 질의응답 가능성과 추적 가능성 88
    • 4.3. 한계 및 향후 연구 방향 89
    • Ⅵ 결론 90
    • 참고문헌 94
    • Abstract 104
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼