Korea’s total R&D costs in 2014 were as high as KRW 63.7 trillion with 7.5% increase compared to those of previous year. In terms of R&D investments of the gross domestic product (GDP), Korea is the highest among OECD states with 4.29%. In a...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T14351533
대전 : 배재대학교, 2017
학위논문(박사) -- 배재대학교 대학원 , 컴퓨터공학과 컴퓨터공학전공 , 2017
2017
한국어
005.74 판사항(6)
005.74 판사항(23)
대전
x, 107장 : 삽화, 도표 ; 26 cm
지도교수: 정회경
권말부록: 연구개발보고서 XML DTD
XML은 "eXtended Markup Language"의 약어임
참고문헌: 장 73-78
0
상세조회0
다운로드다국어 초록 (Multilingual Abstract)
Korea’s total R&D costs in 2014 were as high as KRW 63.7 trillion with 7.5% increase compared to those of previous year. In terms of R&D investments of the gross domestic product (GDP), Korea is the highest among OECD states with 4.29%. In a...
Korea’s total R&D costs in 2014 were as high as KRW 63.7 trillion with 7.5% increase compared to those of previous year. In terms of R&D investments of the gross domestic product (GDP), Korea is the highest among OECD states with 4.29%. In addition, Korean government plans to increase the budget for government-led R&D programs from KRW 18.9 trillion in 2015 to KRW 20.2 trillion by 2019. Due to the continued expansion of the R&D budget, the improvement of investment efficiency in government-led R&D programs and utilization of their outcome have become more important. For this,「ACT ON PERFORMANCE EVALUATION AND MANAGEMENT OF NATIONAL RESEARCH AND DEVELOPMENT PROJECTS, ETC.」 has been enacted and executed since 2006. For the 9 research outcomes resulting from national R&D programs, professional agencies have been designated and operated by research outcome to improve their management efficiency and utilization. The national R&D report, one of the designated research outcomes, is particularly important in that it handles the latest research information and is available as gray literature among researchers. In terms of utilization, however, most domestic and foreign services just provide keyword search services only. Therefore, there is a necessity of service advancement. To increase the utilization of research reports fundamentally, furthermore, it is also needed to accomplish qualitative perfection.
There have long been a lot of efforts to emphasize research ethics and improve the quality of the works such as academic papers and college reports, using a plagiarism detection system. However, the national R&D reports which have been developed with a lot of money have been mostly published without particular procedures, causing several problems.
Therefore, this study proposed a plan to develop an XML DTD by referring to the reports of diverse national R&D programs which were performed by 17 government departments and officers to improve the utilization of national R&D reports and provide services by combing it with a search engine. For the improvement of qualitative performance, furthermore, this study suggested a method to upgrade the quality of research reports by providing national R&D report database and XML element weight-based similarity analysis services when a result report is submitted at the end of a project.
After analyzing about 83,000 national R&D reports in a PDF format in the research outcome management institute ‘KISTI,’ their structure was upgraded according to the National Library of Medicine(NLM) Standard DTD. Then, database was developed by indexing by element, using a search engine. Then, a method to find a wanted sentence in a phrase or chapter more easily than the conventional page-based full-text search service was proposed. Using non-text (tables, figures) information in a research report, moreover, non-texts were automatically extracted. Then, a method to provide a search service on these non-texts was suggested as well. The accuracy of similarity research was improved by dividing the XML research reports by an XML element and imposing weights. Since research reports are as large as 150 pages in average, their similarity was measured in two steps in order: fingerprinting and term appearance pattern analysis. First, the fingerprinting method is advantageous in its fast process without any influence by the length of documents such as research reports by calculating similarity using statistical numbers such as vocabulary, special character, blank and token. However, accuracy tends to decline when many same words appear. To make up for this problem, the term appearance pattern-based method was used to measure the similarity of the research reports in which a group of similar candidates was derived using the fingerprinting method.
After extracting the metadata of the collected national R&D report, sentences were classified according to the sentence segmentation algorithm. Using the fingerprinting function, then, indexing database for the fixed integer was constructed in an inverted index file. Then, the similarity of each sentence was calculated, using the term appearance pattern method. After that, similarity of all research reports was estimated.
The differences between the research report similarity check service and conventional service suggested in this study are as follows: In sentence similarity analysis which has been used in previous studies using a search engine, similarity is analyzed by calculating the number of similar sentences against the total number of sentences in comparison to previous documents after classifying the document into several sentences. However, this study proposes a method to check the similarity of large documents efficiently by covering research reports into XML documents using the e-document XML converter and imposing weights by element. The similarity search service proposed in this study was recognized as a significant similarity analysis service which differs from the conventional services. It appears that this method would enhance researchers’ satisfaction by providing fast and convenient search environment.
It is anticipated that the method mentioned in this study would make a contribution the improvement of the quality of research outcome after being easily applied to the document management system in diverse fields, used by research management institutes and research outcome management agencies which manage large documents such as national R&D reports and provide related services. To increase the efficiency of national R&D programs, furthermore, it would be able to suggest better results than conventional systems through comparison with other results (e.g., technical summary information and software information resulting from the projects, etc.) when similar and redundant projects are searched. To make this method easily applicable to top 9 research outcomes, then, there should be further studies in a continued and systematic manner.
2014년도 우리나라의 총 연구개발비는 63.7조원으로 전년대비 7.5% 증가하였으며, 국내총생산(GDP) 대비 연구개발비 투자비중은 4.29%로 OECD 국가 중 가장 높은 수준이다. 또한 정부 연구개발예산...
2014년도 우리나라의 총 연구개발비는 63.7조원으로 전년대비 7.5% 증가하였으며, 국내총생산(GDP) 대비 연구개발비 투자비중은 4.29%로 OECD 국가 중 가장 높은 수준이다. 또한 정부 연구개발예산도 2015년 18.9조원에서 2019년까지 20.2조원으로 지출규모를 확대할 계획이다. 이러한 정부 연구개발예산의 지속적인 지출규모 확대에 따라 국가연구개발사업의 투자 효율성 제고와 성과 활용에 대한 중요성이 증대되고 있으며, 이를 위해 2006년부터 「국가연구개발사업 등의 성과평가 및 성과관리에 관한 법률」이 제정되어 시행되고 있다. 이를 근거로 국가연구개발사업을 통해 발생한 9개의 연구 성과물의 관리 효율성과 활용도를 높이기 위하여 연구 성과물 별로 전담기관을 지정하여 운영하고 있다.
지정된 연구 성과물 중의 하나인 보고서원문 성과물은 최신 연구정보를 다루고 있고, 연구자들이 회색문헌 중에서 대표적으로 활용하는 학술정보원으로서 중요성을 갖고 있다. 그러나 활용 측면에서는 국내․외 대부분의 서비스는 단순한 키워드 검색서비스를 제공하는 수준에 머물고 있어 서비스 고도화가 필요하다. 더 나아가 연구개발보고서의 활용성을 근본적으로 높이기 위해서는 질적인 완성도를 높이는 방안도 요구되고 있다.
현재 학술논문, 대학의 리포트를 포함한 많은 저작물들에 대해서는 표절 검사시스템을 이용하여 연구윤리 강조 및 저작물의 질적 제고를 오래전부터 강화하여 실행하고 있지만, 많은 비용이 들어간 연구개발보고서에 대해서는 대부분 별도의 절차나 장치 없이 발간되고 관리됨으로써 질적인 측면에서 여러 가지 문제점들을 갖고 있다.
이에 본 논문에서는 연구개발보고서의 활용성을 높이기 위해 17개 부처․청에서 수행한 다양한 국가연구개발사업의 보고서원문을 참고하여 NLM 표준 DTD에 맞게 구조를 개선하였고, 이를 검색엔진과 결합하여 서비스하는 방안을 제안하였다. 연구개발보고서는 평균 150 페이지가 넘는 대용량 문서이기 때문에 연구개발보고서의 유사도를 검사할 때 지문법(fingerprint)과 단어 출현패턴 방식을 차례로 적용하는 2단계로 유사도를 측정하였다. 첫 단계에 사용한 지문법은 단어 혹은 특수문자, 공백, 토큰 등의 통계적 수치를 이용하여 유사도를 계산하므로 연구개발보고서와 같이 커다란 문서의 길이에 영향을 받지 않고 수행속도가 빠르다는 장점이 있지만 동일한 단어들이 많이 등장했을 때 유사도의 정확성이 떨어지는 문제가 있다. 이의 보완을 위해 두 번째는 지문법을 이용하여 유사 후보군을 도출한 연구개발보고서에 대하여 단어 출현패턴 기반 방법으로 유사도를 측정하였다.
또한 연구개발보고서의 질적인 완성도를 높이기 위한 방안으로 과제 종료 후 연구결과 보고서를 제출할 때 선행 국가연구개발과제의 연구개발보고서 DB와 비교하여 XML 엘리먼트 가중치 기반 유사도 분석서비스를 제공함으로써 연구개발보고서의 질적 수준을 제고할 수 있는 방안을 제안하였다.
우선 수집된 연구개발보고서에서 메타데이터를 추출하고, 문장 분리 알고리즘에 따라 문장들로 구분하고 지문법 함수를 이용하여 고정된 정수 값의 색인 DB를 구성하고, 이를 역색인 파일로 구성하였다. 이 후 각각의 문장은 단어 출현패턴 방법에 따라 유사도를 계산하여 모든 연구개발보고서간의 유사도를 계산하였다.
본 논문에서 제시한 연구개발보고서 유사도 검사서비스와 기존 서비스와의 차이점을 살펴보면 다음과 같다. 표절 검색엔진을 비롯한 기존 연구에서 사용하는 문장단위 유사도 분석은 단순히 문서 전체를 일정크기의 문장으로 구분하여 기존의 문서와 비교하여 전체문장 개수 대비 유사한 문장의 개수를 계산하여 유사도를 분석하고 있다. 본 논문에서는 연구개발보고서를 전자문서 XML변환기를 이용하여 XML 문서로 변환하고, 장(Chapter) 유형별로 가중치를 부여할 수 있게 함으로써 대용량 문서를 대상으로 효율적으로 유사도 검사를 할 수 있도록 개선하여 실험하였다. 실험 결과 연구개발보고서를 대표할 수 있는 장 유형은 본문임을 확인하였고, 가중치를 부여하거나 타 유형의 장의 가중치를 낮추는 방법으로 유사도 정확도를 높일 수 있는 방법을 증명하였다.
본 논문에서 제시한 방법을 연구개발보고서를 비롯한 학술정보 등 대용량의 문서를 대상으로 관리 및 서비스를 수행하는 과제관리전문기관, 성과물 전담기관 등에서 사용하는 다양한 분야의 문서관리시스템에 쉽게 적용함으로써 연구 성과물의 질적 제고에 기여할 것으로 기대된다. 또한 국가연구개발사업의 효율성을 높이기 위하여 유사․중복과제를 검색할 때 기존의 과제명, 키워드, 요약 등 메타항목 이외에 과제로부터 발생한 연구개발보고서와도 비교함으로써 기존 시스템보다 정확한 유사도를 제시할 수 있을 것이다.
향후에도 지속적이고 체계적인 연구를 수행하여 9대 연구 성과물을 대상으로 쉽게 적용할 수 있도록 연구가 진행되어야 할 것으로 판단된다.
목차 (Table of Contents)
참고문헌 (Reference)
2. “패턴매칭을 이용한 유사도 비교 분석,”, 고방원, 김영철, 한국컴퓨터정보학회, 한국컴퓨터정보학회 논문지, Vol. 15, No. 1, pp. 187-192, , 2010
3. 특허정보를 활용한 R&D 과제 유사도 측정 모델, 김종배, 김융, 김태균, 선동주, 변정원, 한국정보통신학회논문지, Vol. 18, No. 5, pp. 1013-1021, , 2014
4. 『공공연구기관 연구성과 관리현황 실태조사』, 임채윤, 손수정, 유헌종, 과학기술정책연구원, 과학 기술정책연구원, , 2011
6. “국가연구개발사업의 공공가치 개념 도입방안,”, 조현대, 정윤성, 서지영, 윤문섭, 김명관, 과학기술정책연구원, 과학기술정책연구원, , 2015
9. “대용량 한글 문서를 위한 표절 검색시스템 개발,”, 전명재, 석사학위논문, 부산대학교 컴퓨터공학과, , 2005
2. “패턴매칭을 이용한 유사도 비교 분석,”, 고방원, 김영철, 한국컴퓨터정보학회, 한국컴퓨터정보학회 논문지, Vol. 15, No. 1, pp. 187-192, , 2010
3. 특허정보를 활용한 R&D 과제 유사도 측정 모델, 김종배, 김융, 김태균, 선동주, 변정원, 한국정보통신학회논문지, Vol. 18, No. 5, pp. 1013-1021, , 2014
4. 『공공연구기관 연구성과 관리현황 실태조사』, 임채윤, 손수정, 유헌종, 과학기술정책연구원, 과학 기술정책연구원, , 2011
6. “국가연구개발사업의 공공가치 개념 도입방안,”, 조현대, 정윤성, 서지영, 윤문섭, 김명관, 과학기술정책연구원, 과학기술정책연구원, , 2015
9. “대용량 한글 문서를 위한 표절 검색시스템 개발,”, 전명재, 석사학위논문, 부산대학교 컴퓨터공학과, , 2005
12. “문헌의 문장유사도분석 모델에 관한 실증적 연구,”, 이수한, 숭실대학교 대 학원 IT정책경영학과 박사학위논문, , 2012
13. “태그 간 동의어 집합을 통한 XML 문서 유사도 측정,”, 이강석, 김응모, 송인상, 년도 정보과학회 가을 학술발표논문집, Vol. 34, No. 2(C), pp. 29-34, 2007, , 2007
14. 국가R&D보고서 원문의 활용성과 및 경제적 기여도 분석, 최재황, 곽승진, 한국문헌정보학회지 Vol. 49, No. 1, pp. 113-128, , 2015
15. “군 관련 연구용역 78.5%가 표절 ‘의심’ ‘위험’ 등급,”, 최우석, 월간조선, 년 11월호, , 2014
16. “엘리먼트 기반 XML 문서검색의 성능에 관한 실험적 연구,”, 윤소영, 박사 학위논문, 연세대학교 문헌정보학과, , 2005
19. “과학기술분야 연구개발성과의 활용가치 향상에 관한 연구,”, 오건택, 중앙대학교 대학원, 박사학위논문, 중 앙대학교, 문헌정보학과, , 2011
22. “국가연구개발사업 유사․중복 검색시스템 개발을 위한 실증연구,”, 홍세호, 한국과학기술기획평가원, , 2013
24. “한글 블로그 문서에서의 유사중복문서 검색을 위한 SpotGigs 변형 기법,”, 정시앙, 연세대학교 공학대학원 컴퓨터공학 석사학위 논문, , 2010