특허는 개발된 기술에 대한 상세한 정보를 다양한 형태로 포함하고 있으며, 최근에는 누적된 출원 건수 등의 이유로 특허 빅데이터라는 개념이 등장하였다. 또한, 특허 빅데이터를 분석하여...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T16080712
청주 : 청주대학교 대학원, 2022
2022
한국어
005.76 판사항(5)
충청북도
35p. : 삽화; 26cm.
청주대학교 학위논문은 저작권에 의해 보호받습니다.
Development of Method Patent Big Data Visualization for Business Insight
지도교수:박상성
참고문헌수록
I804:43007-200000577607
0
상세조회0
다운로드특허는 개발된 기술에 대한 상세한 정보를 다양한 형태로 포함하고 있으며, 최근에는 누적된 출원 건수 등의 이유로 특허 빅데이터라는 개념이 등장하였다. 또한, 특허 빅데이터를 분석하여...
특허는 개발된 기술에 대한 상세한 정보를 다양한 형태로 포함하고 있으며, 최근에는 누적된 출원 건수 등의 이유로 특허 빅데이터라는 개념이 등장하였다. 또한, 특허 빅데이터를 분석하여 다양한 가치를 창출하고 경영정보로써, 활용하는 사례가 증가하고 있다. 특허 빅데이터는 분석의 효율성과 결과의 재현 가능성 제고를 위해 통계 및 기계학습 알고리즘을 사용하는 정량적 분석 ( Quantitative Analysis ) 을 주로 수행한다. 또한, 특허 빅데이터 분석자는 정량적 분석의 결과를 효과적으로 전달하기 위해 다양한 시각화 기법 ( Visualization Method ) 을 활용한다.
기존에 특허 빅데이터 시각화는 주요 시장국, 출원인을 식별하기 위해 막대그래프, 파이차트 등을 주로 사용한다. 이와 같은 기법들은 데이터 간의 상대적 비교를 통해 결과 식별에 유용하다. 그러나 비교 대상 데이터가 증가하거나 전체적인 기술 흐름을 파악하고자 할 때는 한계점이 존재한다. 이를 보완하기 위해 일부 연구들에서는 데이터 흐름의 양을 비율적으로 표현해주는 생키 다이어그램 ( Sankey Diagram ) 을 활용하였으나, 세부기술의 식별과 기술 간의 관계를 고려하지 못한다.
본 연구에서는 수집된 특허 빅데이터에 생키 다이어그램과 그래프 모형 ( Graph Model ) 을 사용하여 효과적으로 세부기술을 식별하고 기술개발 흐름 및 기술 간 관계를 시각화하는 방법을 제안한다. 실험 데이터로는 헬스케어 ( Healthcare ) 기술과 관련된 것으로 한국 ( KR ) , 미국 ( US ) , 일본 ( JP ) , 유럽 ( EP ) 에서 출원된 특허를 수집 및 사용한다. 제안하는 방법은 다음과 같다. 먼저, 특허 요약부를 추출하여 텍스트마이닝 ( Text Mining ) 으로 전처리하고 Doc2Vec을 통해 문서를 임베딩한다. 다음으로, 군집화 ( Clustering ) 와 Textrank를 통해 세부기술 ( Elementary Technology ) 을 식별한다. 또한, 정량정보를 추출하여 생키 다이어그램으로 기술개발 흐름을 시각화한다. 마지막으로, 상관행렬의 구축 및 그래프 모형을 통해 기술 간 관계를 식별한다.
실험결과, 헬스케어 기술을 4개의 세부기술 군집으로 식별할 수 있었으며 Textrank를 통해 세부기술 군집별 영향력이 높은 단어들을 산출하고 기술 정의를 수행하였다. 또한, 식별된 세부기술과 특허 정량정보를 결합하여 생키 다이어그램으로 표현하고 최근 스마트폰 연동 생체계측 기술이 주로 개발되고 있는 것을 확인하였다. 추가적으로 기술 간 영향 정도를 파악하기 위해 기술 관계 그래프를 도출하고 통신 관련 기술이 헬스케어 기술 분야에서 핵심적인 역할을 하는 것으로 도출하였다.
다국어 초록 (Multilingual Abstract)
Patents contain detailed information about the developed technology in various forms. Recently, the concept of a patent big data has emerged for reasons such as the number of accumulated applications. In addition, the cases of creating various values ...
Patents contain detailed information about the developed technology in various forms. Recently, the concept of a patent big data has emerged for reasons such as the number of accumulated applications. In addition, the cases of creating various values by analyzing the patent big data and using it as management information are increasing. The Patent big data mainly performs quantitative analysis using statistics and machine learning algorithms to improve analysis efficiency and reproducibility of results. Moreover, the patent big data analyst utilizes various visualization methods to effectively deliver the results of quantitative analysis.
In the past, patent big data visualization mainly uses bar graphs and pie charts to identify major market countries and applicants. Such visualization methods are useful for identifying results through relative comparison between data. However, there are limitations when the data to be compared increases or when trying to understand the overall technology flow. To solve this problem, some studies used a sankey diagram, which expresses the amount of data flow proportionally, but it does not consider the relationship between the identification of detailed technology and the technology.
This paper proposes a method for effectively identifying detailed technologies and visualizing the technology development flow and the relationship between technologies using sankey diagrams and graph models for patent big data. As experimental data, patents applied in Korea(KR), US(US), Japan(JP), and Europe(EP) are collected and used as it relates to healthcare technology. The suggested method is as follows. First, the patent abstracts is extracted, pre-processed by applying text mining, and the document is embedded through Doc2Vec. Next, detailed technologies are identified through clustering and textrank. Furthermore, by extracting quantitative information, the flow of technology development is visualized with the sankey diagram. Finally, the relationship between technologies is identified through the construction of a correlation matrix and the graph model.
As a result of the experiment, healthcare technology could be identified as four elementary technology clusters, and words with high influence for each elementary technology cluster were calculated through textrank and the technology definition was performed. Moreover, it was confirmed that the identified detailed technology and patent quantitative information were combined and expressed as the sankey diagram, and that biometrics technology linked to smartphones is being mainly developed. In addition, a technology relationship graph was derived to understand the degree of influence between technologies, and it was derived that communication-related technologies play a key role in the healthcare technology field.
목차 (Table of Contents)