RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    멀티 플랫폼 데이터를 활용한 문서 유사도 기반 리포트 생성 시스템 = Report Generation System Based on Document Similarity Using Multi-Platform Data

    한글로보기

    https://www.riss.kr/link?id=T17405606

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 연구는 대량의 텍스트 문서를 효율적으로 분석하고 의미 있는 인사이트를 자동으로 생성하는 AI 기반 문서 유사도 리포트 시스템을 설계·구현하였다. 정보 과부하의 시대에 사용자는 방대한 문서에 노출되고 있으나, 이들을 수작업으로 분석하는 것은 시간이 많이 소요되고 비효율적이다. 특히 유사한 내용을 다루는 문서들을 자동으로 탐지하고, 공통점과 차이점을 비교하며, 주제별로 분류하는 작업은 현대 정보 관리의 핵심 과제이다.
    본 연구에서 구현한 시스템은 Spring Boot(3.x), MySQL(9.x), OpenAI API 등 최신 기술 스택 등 최신 기술 스택을 기반으로 개발되었다. 텍스트 문서를 OpenAI Embeddings API를 통해 1536차원의 벡터로 변환한 뒤 코사인 유사도를 적용해 문서 간 유사도를 계산한다. 이후 GPT-4o의 자연어 생성 능력을 활용해 ①키워드 기반 유사 문서 추천, ②다중 문서 비교 분석, ③문서 자동 클러스터링 세 가지 유형의 리포트를 자동 생성한다. 관리자와 일반 사용자를 구분한 역할 기반 웹 대시보드를 제공해 다양한 사용자의 요구를 충족한다.
    실험 결과, 평균 22%의 유사도 점수가 측정되었음에도 상위 유사 문서의 80% 이상이 동일 카테고리에 속하였으며, 자동 생성된 클러스터는‘보안’ 및 ‘기후’ 주제에서 의미 있는 하위 분류를 보여주었다. 이는 절대 유사도보다 상대적 순위 기반 평가가 더 유의미함을 시사한다. GPT-4o의 자연어 생성 능력이 낮은 임베딩 유사도의 한계를 보완함으로써, 사용자가 대량의 문서를 신속하게 이해하고 인사이트를 도출할 수 있게 하였다.
    본 연구는 문서 분석 및 정보 시스템의 자동화 가능성을 실증함으로써, 향후 AI 기반 정보 처리 시스템 발전에 기초를 제공한다. 구현된 시스템은 기업의 시장 조사, 학술 기관의 연구 동향 분석, 정부 기관의 정책 분석, 언론사의 기사 분류 등 다양한 산업 분야에 응용 가능하다.
    번역하기

    본 연구는 대량의 텍스트 문서를 효율적으로 분석하고 의미 있는 인사이트를 자동으로 생성하는 AI 기반 문서 유사도 리포트 시스템을 설계·구현하였다. 정보 과부하의 시대에 사용자는 방...

    본 연구는 대량의 텍스트 문서를 효율적으로 분석하고 의미 있는 인사이트를 자동으로 생성하는 AI 기반 문서 유사도 리포트 시스템을 설계·구현하였다. 정보 과부하의 시대에 사용자는 방대한 문서에 노출되고 있으나, 이들을 수작업으로 분석하는 것은 시간이 많이 소요되고 비효율적이다. 특히 유사한 내용을 다루는 문서들을 자동으로 탐지하고, 공통점과 차이점을 비교하며, 주제별로 분류하는 작업은 현대 정보 관리의 핵심 과제이다.
    본 연구에서 구현한 시스템은 Spring Boot(3.x), MySQL(9.x), OpenAI API 등 최신 기술 스택 등 최신 기술 스택을 기반으로 개발되었다. 텍스트 문서를 OpenAI Embeddings API를 통해 1536차원의 벡터로 변환한 뒤 코사인 유사도를 적용해 문서 간 유사도를 계산한다. 이후 GPT-4o의 자연어 생성 능력을 활용해 ①키워드 기반 유사 문서 추천, ②다중 문서 비교 분석, ③문서 자동 클러스터링 세 가지 유형의 리포트를 자동 생성한다. 관리자와 일반 사용자를 구분한 역할 기반 웹 대시보드를 제공해 다양한 사용자의 요구를 충족한다.
    실험 결과, 평균 22%의 유사도 점수가 측정되었음에도 상위 유사 문서의 80% 이상이 동일 카테고리에 속하였으며, 자동 생성된 클러스터는‘보안’ 및 ‘기후’ 주제에서 의미 있는 하위 분류를 보여주었다. 이는 절대 유사도보다 상대적 순위 기반 평가가 더 유의미함을 시사한다. GPT-4o의 자연어 생성 능력이 낮은 임베딩 유사도의 한계를 보완함으로써, 사용자가 대량의 문서를 신속하게 이해하고 인사이트를 도출할 수 있게 하였다.
    본 연구는 문서 분석 및 정보 시스템의 자동화 가능성을 실증함으로써, 향후 AI 기반 정보 처리 시스템 발전에 기초를 제공한다. 구현된 시스템은 기업의 시장 조사, 학술 기관의 연구 동향 분석, 정부 기관의 정책 분석, 언론사의 기사 분류 등 다양한 산업 분야에 응용 가능하다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study designed and implemented an AI-based document similarity report system that efficiently analyzes a large amount of text documents and automatically generates meaningful insights. In an era of information overload, users are exposed to vast amounts of documents, but analyzing them manually is time-consuming and inefficient. In particular, the task of automatically detecting documents dealing with similar content, comparing commonalities and differences, and classifying them by subject is a key task of modern information management.
    The system implemented in this study was developed based on the latest technology stack such as Spring Boot (3.x), MySQL (9.x), and OpenAI API. After converting text documents into 1536-dimensional vectors through the OpenAI Embedding API, cosine similarity is applied to calculate the similarity between documents. Afterwards, GPT-4o's natural language generation capabilities are utilized to automatically generate three types of reports: ① keyword-based similar documents recommendation, ② multi-document comparison analysis, and ③ automatic document clustering. It meets the needs of various users by providing a role-based web dashboard that distinguishes administrators and general users.
    As a result of the experiment, even though an average similarity score of 22% was measured, more than 80% of the top similar documents belonged to the same category, and the automatically generated cluster showed meaningful subclassification in the topic of 'security' and 'climate'. This suggests that the relative ranking-based evaluation is more significant than the absolute similarity. By supplementing the limit of embedding similarity, GPT-4o's natural language generation ability is low, it enabled users to quickly understand a large amount of documents and derive insights.
    This study provides the basis for the development of AI-based information processing systems in the future by demonstrating the possibility of automation of document analysis and information systems. The implemented system can be applied to various industrial fields such as corporate market research, analysis of research trends of academic institutions, policy analysis of government agencies, and classification of articles of media companies.
    번역하기

    This study designed and implemented an AI-based document similarity report system that efficiently analyzes a large amount of text documents and automatically generates meaningful insights. In an era of information overload, users are exposed to vast ...

    This study designed and implemented an AI-based document similarity report system that efficiently analyzes a large amount of text documents and automatically generates meaningful insights. In an era of information overload, users are exposed to vast amounts of documents, but analyzing them manually is time-consuming and inefficient. In particular, the task of automatically detecting documents dealing with similar content, comparing commonalities and differences, and classifying them by subject is a key task of modern information management.
    The system implemented in this study was developed based on the latest technology stack such as Spring Boot (3.x), MySQL (9.x), and OpenAI API. After converting text documents into 1536-dimensional vectors through the OpenAI Embedding API, cosine similarity is applied to calculate the similarity between documents. Afterwards, GPT-4o's natural language generation capabilities are utilized to automatically generate three types of reports: ① keyword-based similar documents recommendation, ② multi-document comparison analysis, and ③ automatic document clustering. It meets the needs of various users by providing a role-based web dashboard that distinguishes administrators and general users.
    As a result of the experiment, even though an average similarity score of 22% was measured, more than 80% of the top similar documents belonged to the same category, and the automatically generated cluster showed meaningful subclassification in the topic of 'security' and 'climate'. This suggests that the relative ranking-based evaluation is more significant than the absolute similarity. By supplementing the limit of embedding similarity, GPT-4o's natural language generation ability is low, it enabled users to quickly understand a large amount of documents and derive insights.
    This study provides the basis for the development of AI-based information processing systems in the future by demonstrating the possibility of automation of document analysis and information systems. The implemented system can be applied to various industrial fields such as corporate market research, analysis of research trends of academic institutions, policy analysis of government agencies, and classification of articles of media companies.

    더보기

    목차 (Table of Contents)

    • 국문초록ⅰ
    • 목 차ⅲ
    • 그림목차ⅴ
    • 도표목차ⅶ
    • I. 서 론 1
    • 국문초록ⅰ
    • 목 차ⅲ
    • 그림목차ⅴ
    • 도표목차ⅶ
    • I. 서 론 1
    • 1.1 연구의 배경 및 필요성 1
    • 1.2 연구 목적 3
    • 1.3 연구 범위 및 구성5
    • Ⅱ. 관련 연구 및 이론적 배경 6
    • 2.1 문서 분석 및 정보 추출 관련 기술6
    • 2.1.1 텍스트 수집 및 전처리 기법7
    • 2.1.2 임베딩 기반 문서 벡터화8
    • 2.1.3 유사도 계산 및 클러스터링10
    • 2.2 자동 요약 및 생성형 AI 활용12
    • 2.2.1 GPT 계열 모델의 적용 사례 12
    • 2.2.2 인공지능 기반 리포팅의 동향 14
    • 2.3 데이터 시각화 및 대시보드 설계14
    • 2.4 연구의 차별성과 기존 연구 한계15
    • Ⅲ. 시스템 설계 및 구현 18
    • 3.1 전체 시스템 아키텍처 및 기술 스택 18
    • 3.1.1 시스템 구조 및 처리 흐름 20
    • 3.1.2 도입 기술 환경 및 구성요소 22
    • 3.2 데이터 수집 시스템 26
    • 3.2.1 클라이언트 사이드 트래킹 구조 26
    • 3.2.2 텍스트 데이터 수집 설계 28
    • 3.2.3 임시 데이터셋 구조화 및 활용30
    • 3.3 문서 임베딩 및 유사도 분석34
    • 3.3.1 OpenAI Embeddings API 연동 34
    • 3.3.2 벡터 데이터 저장 및 관리 36
    • 3.4 자동 리포트 생성 시스템37
    • 3.4.1 GPT-4o 기반 요약 및 키워드 추출 37
    • 3.4.2 키워드 기반 유사 문서 추천 리포트 40
    • 3.4.3 다중 문서 비교 리포트 41
    • 3.4.4 문서 클러스터링 리포트 44
    • 3.5 사용자 인터페이스 및 시각화 47
    • 3.5.1 관리자 대시보드 47
    • 3.5.2 애널리틱스 대시보드 49
    • 3.5.3 텍스트 분석 대시보드50
    • 3.6 데이터베이스 및 테이블 구조 51
    • Ⅳ. 실험 및 결과 분석 59
    • 4.1 실험 환경 및 임시 데이터셋59
    • 4.2 유사도 분석 알고리즘 실험 61
    • 4.3 리포트 자동 생성 결과 64
    • Ⅴ. 결론 71
    • 참고문헌 74
    • 영문초록 78
    • 감사의 글(Acknowledgement) 80
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼