RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    CNN-based handwritten character recognition for the manchu and hangul scripts = CNN 기반 만주 및 한글 필기 스크립트 문자 인식

    한글로보기

    https://www.riss.kr/link?id=T16928583

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    현대 기술은 인간의 능력과 이해를 돕고 대량의 데이터를 신속하게 처리하여 더 쉽게 접근할 수 있도록 설계되었다. 광학 문자 인식(OCR)은 인쇄된 텍스트의 디지털화를 돕는 기술 중 하나이다. 그러나 최고의 OCR 모델조차도 손으로 쓴 텍스트를 정확하게 인식하는 데 여전히 어려움을 겪고 있다. 특히 역사적 문서와 일반적이지 않은 스크립트의 OCR은 더욱 그러하다. 또한, 역사적 문서에서 손으로 쓴 스크립트 문자의 대표 샘플을 대량 기계 학습 데이터 세트를 생성할 수 있을 만큼 충분하고도 다양한 스타일 변형을 얻는 것은 어렵다.

    이 논문은 필기 문자 인식(HCR)을 위해 CNN(Convolutional Neural Networks)에서 훈련할 수 있도록 대량의 다양한 필기 스크립트 문자 데이터 세트를 생성하는 문제를 다룬다. 이 논문에서는 매우 다른 두 가지 아시아 문자인 만주 문자와 한국어 한글의 필기체를 연구하였다. 둘 다 쉽지 않은 문제이다.

    만주 스크립트는 ‘심각한 멸종 위기에 처한’ 만주어를 표현하는 역사적인 알파벳 문자이다. 여러 보고가 있지만 여전히 만주어를 읽고 쓸 수 있는 사람은 수십에서 수백 명에 불과한 것으로 추산된다. 따라서 2023년에 수집할 수 있는 만주 문자 샘플의 대부분은 주로 전 세계 기록 보관소에 흩어져 있는 역사적 문서로 구성되었다.

    반면, 한글은 현재 사용되고 있는 한국의 기본 문자 체계이다. 또한, 한국 대중문화에 대한 관심이 전 세계적으로 확산되면서 손으로 직접 쓴 한글 샘플을 얻기는 어렵지 않다. 따라서 한글 샘플을 수집하는 것은 비교적 쉽다. 그러나 두 스크립트의 서로 다른 쓰기 스타일은 쉽지 않은 양상을 제기한다.

    만주(Manchu) 스크립트는 위에서 아래로 수직으로 쓰여지며 단어의 각 글자는 중앙 어간으로 연결된다. 만주어 문자 사이에는 공백 구분이 없다. 이는 만주어를 세로가 아닌 가로로 쓰는 영어 필기체와 어떤 면에서 유사하게 만든다. 만주어의 개별 문자를 분할하고 감지하는 것은 각 문자가 다음 문자로 흘러가기 때문에 특히 어렵다. 일부 문자 디센더는 다른 문자의 가로 공간과 겹치고 일부 문자 모양은 매우 미묘하다. 이는 문자의 자동 분할을 훨씬 어렵게 만든다.

    그러나 한글 단어는 2차원 음절 블록에서 왼쪽에서 오른쪽으로 가로로 작성된다. 한글 음절 블록은 자음과 모음 문자의 조합으로, 중첩되거나 결합되어 거의 10가지 방법으로 생성된다. 이런 점에서 볼 때 한글 음절 블록은 단어의 모든 문자가 한 방향으로 흐르는 영어나 만주어보다는 정사각형 블록으로 작성되는 전통 중국어 문자 즉, 한자와 더 유사해진다.

    공백을 기반으로 한 음절 블록의 자동 분할은 비교적 쉽다. 그래서 공백을 기준으로 한글 음절의 문자를 분리하는 것도 비교적 쉽다. 그러나 24개의 기본 음소(자음, 모음) 모양은 서로 다른 음절 블록에서 총 11,172개의 음절을 만들어 낸다. 이로 인해 각 음절 블록을 개별적으로 분류하도록 CNN을 훈련시키는 것이 거의 불가능해졌다. 개별 한글 문자를 식별하도록 CNN을 훈련시키는 것이 음절 블록을 사용하여 식별하는 것보다 더 합리적이다. 그러나 음절 블록의 2차원 특성으로 인해 OCR에 추가적인 문제가 발생한다.

    본 논문은 위에 제시된 각각의 문제를 해결하려고 시도한다. 첫쩨, 스캔한 수백 페이지의 만주 스크립트 이미지 중에서 문자 템플릿을 선택하고 일치시켜 만주 문자 데이터를 수집하였다. 둘째, 한국 대학생들을 대상으로 손으로 쓴 한글 자료를 수집하였다. 셋째, 각 스크립트에 대한 두 개의 서로 다른 데이터 세트가 CNN에서 생성, 정규화 및 훈련되었다. 경우에 따라서는 스크립트 문자 데이터의 양이 부족한 경우가 있었다. 따라서 CNN 훈련 데이터를 보완하기 위해 각 데이터 세트마다 ACGAN(Auxiliary Classifier Generative Adversarial Network) 모델도 개발되었다.마지막으로 최종 CNN 모델은 간단한 OCR 시도로 손으로 쓴 텍스트 페이지의 실제 예에 적용되었다. 비록 본 논문에 적용된 OCR 알고리즘이 완벽하지는 않지만 데이터 세트의 약점을 파악하고 향후 연구 기회를 제공하는 데 기여하였다.

    이 논문의 주요 기여는 데이터를 수집하고 데이터 세트를 생성 및 정규화하는 새로운 방법을 제시했다는 데 있다. 본 연구 결과, 89,100개의 이미지로 구성된 한글 전체 데이터 세트가 생성되었으며, 만주 문자의 분할 및 라벨링을 위한 새로운 방법이 개발되었다.
    번역하기

    현대 기술은 인간의 능력과 이해를 돕고 대량의 데이터를 신속하게 처리하여 더 쉽게 접근할 수 있도록 설계되었다. 광학 문자 인식(OCR)은 인쇄된 텍스트의 디지털화를 돕는 기술 중 하나이...

    현대 기술은 인간의 능력과 이해를 돕고 대량의 데이터를 신속하게 처리하여 더 쉽게 접근할 수 있도록 설계되었다. 광학 문자 인식(OCR)은 인쇄된 텍스트의 디지털화를 돕는 기술 중 하나이다. 그러나 최고의 OCR 모델조차도 손으로 쓴 텍스트를 정확하게 인식하는 데 여전히 어려움을 겪고 있다. 특히 역사적 문서와 일반적이지 않은 스크립트의 OCR은 더욱 그러하다. 또한, 역사적 문서에서 손으로 쓴 스크립트 문자의 대표 샘플을 대량 기계 학습 데이터 세트를 생성할 수 있을 만큼 충분하고도 다양한 스타일 변형을 얻는 것은 어렵다.

    이 논문은 필기 문자 인식(HCR)을 위해 CNN(Convolutional Neural Networks)에서 훈련할 수 있도록 대량의 다양한 필기 스크립트 문자 데이터 세트를 생성하는 문제를 다룬다. 이 논문에서는 매우 다른 두 가지 아시아 문자인 만주 문자와 한국어 한글의 필기체를 연구하였다. 둘 다 쉽지 않은 문제이다.

    만주 스크립트는 ‘심각한 멸종 위기에 처한’ 만주어를 표현하는 역사적인 알파벳 문자이다. 여러 보고가 있지만 여전히 만주어를 읽고 쓸 수 있는 사람은 수십에서 수백 명에 불과한 것으로 추산된다. 따라서 2023년에 수집할 수 있는 만주 문자 샘플의 대부분은 주로 전 세계 기록 보관소에 흩어져 있는 역사적 문서로 구성되었다.

    반면, 한글은 현재 사용되고 있는 한국의 기본 문자 체계이다. 또한, 한국 대중문화에 대한 관심이 전 세계적으로 확산되면서 손으로 직접 쓴 한글 샘플을 얻기는 어렵지 않다. 따라서 한글 샘플을 수집하는 것은 비교적 쉽다. 그러나 두 스크립트의 서로 다른 쓰기 스타일은 쉽지 않은 양상을 제기한다.

    만주(Manchu) 스크립트는 위에서 아래로 수직으로 쓰여지며 단어의 각 글자는 중앙 어간으로 연결된다. 만주어 문자 사이에는 공백 구분이 없다. 이는 만주어를 세로가 아닌 가로로 쓰는 영어 필기체와 어떤 면에서 유사하게 만든다. 만주어의 개별 문자를 분할하고 감지하는 것은 각 문자가 다음 문자로 흘러가기 때문에 특히 어렵다. 일부 문자 디센더는 다른 문자의 가로 공간과 겹치고 일부 문자 모양은 매우 미묘하다. 이는 문자의 자동 분할을 훨씬 어렵게 만든다.

    그러나 한글 단어는 2차원 음절 블록에서 왼쪽에서 오른쪽으로 가로로 작성된다. 한글 음절 블록은 자음과 모음 문자의 조합으로, 중첩되거나 결합되어 거의 10가지 방법으로 생성된다. 이런 점에서 볼 때 한글 음절 블록은 단어의 모든 문자가 한 방향으로 흐르는 영어나 만주어보다는 정사각형 블록으로 작성되는 전통 중국어 문자 즉, 한자와 더 유사해진다.

    공백을 기반으로 한 음절 블록의 자동 분할은 비교적 쉽다. 그래서 공백을 기준으로 한글 음절의 문자를 분리하는 것도 비교적 쉽다. 그러나 24개의 기본 음소(자음, 모음) 모양은 서로 다른 음절 블록에서 총 11,172개의 음절을 만들어 낸다. 이로 인해 각 음절 블록을 개별적으로 분류하도록 CNN을 훈련시키는 것이 거의 불가능해졌다. 개별 한글 문자를 식별하도록 CNN을 훈련시키는 것이 음절 블록을 사용하여 식별하는 것보다 더 합리적이다. 그러나 음절 블록의 2차원 특성으로 인해 OCR에 추가적인 문제가 발생한다.

    본 논문은 위에 제시된 각각의 문제를 해결하려고 시도한다. 첫쩨, 스캔한 수백 페이지의 만주 스크립트 이미지 중에서 문자 템플릿을 선택하고 일치시켜 만주 문자 데이터를 수집하였다. 둘째, 한국 대학생들을 대상으로 손으로 쓴 한글 자료를 수집하였다. 셋째, 각 스크립트에 대한 두 개의 서로 다른 데이터 세트가 CNN에서 생성, 정규화 및 훈련되었다. 경우에 따라서는 스크립트 문자 데이터의 양이 부족한 경우가 있었다. 따라서 CNN 훈련 데이터를 보완하기 위해 각 데이터 세트마다 ACGAN(Auxiliary Classifier Generative Adversarial Network) 모델도 개발되었다.마지막으로 최종 CNN 모델은 간단한 OCR 시도로 손으로 쓴 텍스트 페이지의 실제 예에 적용되었다. 비록 본 논문에 적용된 OCR 알고리즘이 완벽하지는 않지만 데이터 세트의 약점을 파악하고 향후 연구 기회를 제공하는 데 기여하였다.

    이 논문의 주요 기여는 데이터를 수집하고 데이터 세트를 생성 및 정규화하는 새로운 방법을 제시했다는 데 있다. 본 연구 결과, 89,100개의 이미지로 구성된 한글 전체 데이터 세트가 생성되었으며, 만주 문자의 분할 및 라벨링을 위한 새로운 방법이 개발되었다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Modern technologies are designed to aid human abilities and understanding, as well as to quickly process large amounts of data to make it more accessible. Optical Character Recognition (OCR) is one such technology that aids in the digitization of printed text. However, even the best OCR models still struggle to precisely recognize handwritten text. This is particularly true for OCR in historical documents and uncommon scripts. Additionally, it is difficult to obtain a representative sample of handwritten script letters from historical documents in a large enough quantity and with a wide enough variation of styles that a sufficient machine learning dataset can be created.

    This dissertation addresses the problem of creating handwritten script letter datasets of a sufficiently large and varied quantity, to be trained in Convolutional Neural Networks (CNN) for the purpose of Handwritten Character Recognition (HCR). In this dissertation, two vastly different handwritten Asian scripts were studied: the Manchu script and Korean Hangul. Both present unique challenges.

    Manchu is a historic alphabetic writing script that expresses the ‘critically endangered’ Manchu language. Reports vary, but it is estimated that only between a few dozen to a few hundred people can still read and write Manchu. Therefore, the majority of written Manchu script samples that can be collected in 2023 consist primarily of historic documents scattered in archives around the world.

    Hangul, on the other hand, is a featural writing system that is currently used as the primary writing system of the Koreas. Additionally, interest in Korean pop culture has spread throughout the world, so there is no shortage of handwritten Hangul samples. Therefore, collecting samples of Hangul letters is comparatively easy. However, the differing writing styles for both scripts also present unique challenges.

    Manchu is written vertically from top-to-bottom in lines with each letter in a word connected by a central stem. There is no whitespace separation between Manchu letters. This makes Manchu similar in some ways to the English cursive script, which is written horizontally, rather than vertically. Segmenting and detecting individual letters in Manchu is particularly difficult because each letter flows into the next letter. Some letter descenders overlap into the horizontal space of other letters, and some letter shapes are very subtle. This makes automatic segmentation of letters incredibly difficult.

    Hangul words, however, are written horizontally from left-to-right in two-dimensional syllable blocks. Hangul syllable blocks are a combination of consonant and vowel letters that can be doubled or combined in almost ten different ways. This makes Hangul syllable blocks more similar to traditional Chinese letters (which are also written in square blocks) than to English or Manchu where all letters in a word flow in a single direction.

    The automatic segmentation of syllable blocks based on whitespace is comparatively easy. In fact, segmenting letters from Hangul syllables based on whitespace is also comparatively easy. However, even with only 24 basic letter shapes, there are a total of 11,172 letter combinations in different syllable blocks. This makes training a CNN to individually classify each syllable block nearly impossible. Training a CNN to identify individual Hangul letters makes more sense than attempting to do so with syllable blocks. However, the two-dimensional nature of the syllable blocks poses additional challenges for OCR.

    This dissertation attempts to address each of the problems presented above. First, Manchu script data was gathered by selecting and matching letter templates from within hundreds of pages of scanned Manchu script images. Second, handwritten Hangul script data was gathered from Korean university students. Third, two different datasets for each script were created, normalized, and trained in a CNN. In some cases, the quantity of script letter data for a certain class was insufficient. Therefore, an Auxiliary Classifier Generative Adversarial Network (ACGAN) model was also developed for each dataset in order to supplement the CNN training data. Last, the final CNN models were applied to real examples of pages of handwritten text in a simple attempt at OCR. Although the OCR algorithm applied in this dissertation was not perfect, it was useful in highlighting both weaknesses in the datasets and opportunities for future research.

    The major contributions of this dissertation are a new method for gathering data, and creating and normalizing datasets for both Manchu and Hangul handwritten scripts. As a result of this research, a full Hangul dataset of 89,100 images was created, and a new method for the segmentation and labeling of Manchu script letters was developed.
    번역하기

    Modern technologies are designed to aid human abilities and understanding, as well as to quickly process large amounts of data to make it more accessible. Optical Character Recognition (OCR) is one such technology that aids in the digitization of prin...

    Modern technologies are designed to aid human abilities and understanding, as well as to quickly process large amounts of data to make it more accessible. Optical Character Recognition (OCR) is one such technology that aids in the digitization of printed text. However, even the best OCR models still struggle to precisely recognize handwritten text. This is particularly true for OCR in historical documents and uncommon scripts. Additionally, it is difficult to obtain a representative sample of handwritten script letters from historical documents in a large enough quantity and with a wide enough variation of styles that a sufficient machine learning dataset can be created.

    This dissertation addresses the problem of creating handwritten script letter datasets of a sufficiently large and varied quantity, to be trained in Convolutional Neural Networks (CNN) for the purpose of Handwritten Character Recognition (HCR). In this dissertation, two vastly different handwritten Asian scripts were studied: the Manchu script and Korean Hangul. Both present unique challenges.

    Manchu is a historic alphabetic writing script that expresses the ‘critically endangered’ Manchu language. Reports vary, but it is estimated that only between a few dozen to a few hundred people can still read and write Manchu. Therefore, the majority of written Manchu script samples that can be collected in 2023 consist primarily of historic documents scattered in archives around the world.

    Hangul, on the other hand, is a featural writing system that is currently used as the primary writing system of the Koreas. Additionally, interest in Korean pop culture has spread throughout the world, so there is no shortage of handwritten Hangul samples. Therefore, collecting samples of Hangul letters is comparatively easy. However, the differing writing styles for both scripts also present unique challenges.

    Manchu is written vertically from top-to-bottom in lines with each letter in a word connected by a central stem. There is no whitespace separation between Manchu letters. This makes Manchu similar in some ways to the English cursive script, which is written horizontally, rather than vertically. Segmenting and detecting individual letters in Manchu is particularly difficult because each letter flows into the next letter. Some letter descenders overlap into the horizontal space of other letters, and some letter shapes are very subtle. This makes automatic segmentation of letters incredibly difficult.

    Hangul words, however, are written horizontally from left-to-right in two-dimensional syllable blocks. Hangul syllable blocks are a combination of consonant and vowel letters that can be doubled or combined in almost ten different ways. This makes Hangul syllable blocks more similar to traditional Chinese letters (which are also written in square blocks) than to English or Manchu where all letters in a word flow in a single direction.

    The automatic segmentation of syllable blocks based on whitespace is comparatively easy. In fact, segmenting letters from Hangul syllables based on whitespace is also comparatively easy. However, even with only 24 basic letter shapes, there are a total of 11,172 letter combinations in different syllable blocks. This makes training a CNN to individually classify each syllable block nearly impossible. Training a CNN to identify individual Hangul letters makes more sense than attempting to do so with syllable blocks. However, the two-dimensional nature of the syllable blocks poses additional challenges for OCR.

    This dissertation attempts to address each of the problems presented above. First, Manchu script data was gathered by selecting and matching letter templates from within hundreds of pages of scanned Manchu script images. Second, handwritten Hangul script data was gathered from Korean university students. Third, two different datasets for each script were created, normalized, and trained in a CNN. In some cases, the quantity of script letter data for a certain class was insufficient. Therefore, an Auxiliary Classifier Generative Adversarial Network (ACGAN) model was also developed for each dataset in order to supplement the CNN training data. Last, the final CNN models were applied to real examples of pages of handwritten text in a simple attempt at OCR. Although the OCR algorithm applied in this dissertation was not perfect, it was useful in highlighting both weaknesses in the datasets and opportunities for future research.

    The major contributions of this dissertation are a new method for gathering data, and creating and normalizing datasets for both Manchu and Hangul handwritten scripts. As a result of this research, a full Hangul dataset of 89,100 images was created, and a new method for the segmentation and labeling of Manchu script letters was developed.

    더보기

    목차 (Table of Contents)

    • List of Tables v
    • List of Figures vi
    • Acronyms and Glossary ix
    • Abstract x
    • List of Tables v
    • List of Figures vi
    • Acronyms and Glossary ix
    • Abstract x
    • I. Introduction 1
    • 1.1 Background 1
    • 1.2 Research Goals 3
    • 1.3 Motivation 4
    • 1.4 Contributions 5
    • 1.5 Organization of Dissertation 6
    • II. Background 9
    • 2.1 Introduction 9
    • 2.2 Manchu 14
    • 2.2.1 History of the Manchu Script 16
    • 2.2.2 Characteristics of the Manchu Script 17
    • 2.2.3 OCR Challenges with Manchu 22
    • 2.3 Hangul 24
    • 2.3.1 History of Hangul 25
    • 2.3.2 Characteristics of Hangul 26
    • 2.3.3 OCR Challenges with Hangul 31
    • III. Related Research 34
    • 3.1 Introduction 34
    • 3.2 Manchu OCR 34
    • 3.2.1 Manchu OCR Techniques 35
    • 3.2.2 Manchu OCR Datasets 39
    • 3.2.3 Manchu OCR Performance Metrics 40
    • 3.3 Hangul HCR 40
    • 3.3.1 Hangul HCR Techniques 40
    • 3.3.2 Hangul HCR Datasets 42
    • 3.3.3 Hangul HCR Performance Metrics 43
    • IV. Script Dataset Creation 45
    • 4.1 Introduction 45
    • 4.2 Data Sources 47
    • 4.2.1 Manchu: Textbooks 47
    • 4.2.2 Hangul: Handwriting 48
    • 4.3 Data Extraction 56
    • 4.3.1 Manchu 56
    • 4.3.1.1 Algorithmic Segmentation 57
    • 4.3.1.2 Region of Interest Selection and Template Matching 67
    • 4.3.2 Hangul 73
    • 4.3.2.1 Online Image Splitting 73
    • 4.3.2.2 Sudoku Style Box Analysis 77
    • 4.3.2.3 Hangul: Checkbox Table Cell Detection 80
    • 4.4 Data Labeling 84
    • 4.4.1 Manchu: Template Matching 85
    • 4.4.2 Hangul: Algorithmic Redistribution 86
    • 4.5 Data Normalization 87
    • 4.6 Data Augmentation 92
    • 4.6.1 Auxiliary Classifier Generative Adversarial Network 93
    • 4.6.2 Manchu: ACGAN Results 96
    • 4.6.3 Hangul: ACGAN Results 97
    • V. Neural Network Training 100
    • 5.1 Introduction 100
    • 5.2 Comparison of Three Simple Models 101
    • 5.2.1 Multilayer Perceptron (MLP) 102
    • 5.2.2 Convolutional Neural Network (CNN) 104
    • 5.2.3 Recurrent Neural Network (RNN) 106
    • 5.3 Training on a Deeper Architecture (ResNet50) 112
    • VI. Optical Character Recognition 120
    • 6.1 Introduction 120
    • 6.2 Recognizing Letters in Real Images 120
    • 6.2.1 Manchu OCR Predictions 121
    • 6.2.2 Revised Manchu OCR Algorithm Design 124
    • 6.2.3 Hangul OCR Predictions 126
    • 6.2.4 Revised Hangul OCR Algorithm Design 128
    • 6.3 Proposed Postprocessing Method with BOW 129
    • VII. Conclusion 131
    • 7.1 Concluding Remarks 131
    • 7.2 Future Work 134
    • References 136
    • Abstract (in Korean) 144
    더보기

    참고문헌 (Reference)

    1. HangulDB, I. J. Kim, Available https://github. com/callee2006/HangulDB, , 2019

    2. Manchu Alphabet and Language, Omniglot, Available https://www. omniglot. com/writing/manchu. htm, , 2023

    3. OpenCV Sudoku Solver and OCR, A. Rosebrock, blog Available https://pyimagesearch. com/2020/08/10/opencv-sudoku-solver-and-ocr/, , 2020

    4. MNIST handwritten digit database, Y. LeCun, C. Cortes, ATT Labs Available http//yann lecun. com/exdb/mnist/, , 2010

    5. Manchu Language Lives Mostly in Archives, D. Lague, Online Available https//www. nytimes. com/2007/03/17/world/asia/18manchu_side. html, , 2007

    6. For Historians, Loss of Manchu Language Would Be a Blow, D. Lague, Available https//www. nytimes. com/2007/03/16/world/asia/16iht-manchuside.4930293 . html, , 2007

    7. Construction of Printed Hangul Character Database PHD08,, D. S. Ham, D. R. Lee, I. S. Oh, I. S. Jung, vol. 8, no. 11, pp. 33–40 DOI: 10.5392/jkca.2008.8.11.033, , 2008

    8. OCR: Handwriting Recognition with OpenCV, Keras, and TensorFlow, A. Rosebrock, blogOnline Available https://pyimagesearch. com/2020/08/24/ocr-handwriting-recognition-withopencv- keras-and-tensorflow/, , 2020

    9. A Systematic Review of Handwritten Hangul OCR Techniques and Datasets,, A. D. Snowberger, C. H. Lee, vol. 3, pp. 27-38, 2023. DOI: http://dx. doi. org/10.59434/JQoLR.2023.1.2.027, , 2023

    10. Handwritten Hangul Recognition Model Using Multi-Label Classification,, H. Choi, vol. 27, no. 2, pp. 135-45, 2023, , 2023

    1. HangulDB, I. J. Kim, Available https://github. com/callee2006/HangulDB, , 2019

    2. Manchu Alphabet and Language, Omniglot, Available https://www. omniglot. com/writing/manchu. htm, , 2023

    3. OpenCV Sudoku Solver and OCR, A. Rosebrock, blog Available https://pyimagesearch. com/2020/08/10/opencv-sudoku-solver-and-ocr/, , 2020

    4. MNIST handwritten digit database, Y. LeCun, C. Cortes, ATT Labs Available http//yann lecun. com/exdb/mnist/, , 2010

    5. Manchu Language Lives Mostly in Archives, D. Lague, Online Available https//www. nytimes. com/2007/03/17/world/asia/18manchu_side. html, , 2007

    6. For Historians, Loss of Manchu Language Would Be a Blow, D. Lague, Available https//www. nytimes. com/2007/03/16/world/asia/16iht-manchuside.4930293 . html, , 2007

    7. Construction of Printed Hangul Character Database PHD08,, D. S. Ham, D. R. Lee, I. S. Oh, I. S. Jung, vol. 8, no. 11, pp. 33–40 DOI: 10.5392/jkca.2008.8.11.033, , 2008

    8. OCR: Handwriting Recognition with OpenCV, Keras, and TensorFlow, A. Rosebrock, blogOnline Available https://pyimagesearch. com/2020/08/24/ocr-handwriting-recognition-withopencv- keras-and-tensorflow/, , 2020

    9. A Systematic Review of Handwritten Hangul OCR Techniques and Datasets,, A. D. Snowberger, C. H. Lee, vol. 3, pp. 27-38, 2023. DOI: http://dx. doi. org/10.59434/JQoLR.2023.1.2.027, , 2023

    10. Handwritten Hangul Recognition Model Using Multi-Label Classification,, H. Choi, vol. 27, no. 2, pp. 135-45, 2023, , 2023

    11. Handwritten Hangul Graphemes Classification Using Three Artificial Neural Networks,, A. D. Snowberger, C. H. Lee, vol. 21, no. 2, pp. 167-173, 2023. DOI: https://doi. org/10.56977/jicce.2023.21.2.167, , 2023

    12. Synthetic Data and DAG-SVM Classifier for Segmentation-Free Manchu Word Recognition,, J. Bi, S. Xu, D. Huang, R. Zheng, M. Li, 2017 International Conference on Computing Intelligence and Information System (CIIS), Nanjing, China, pp. 46-50, 2017. DOI: 10.1109/CIIS.2017.15, , 2017

    13. An offline recognition method of handwritten primitive Manchu characters based on strokes, R. W. He, J. J. Li, G. Y. Zhang, A. X. Wang, Ninth International Workshop on Frontiers in Handwriting Recognition, Kokubunji, Japan, pp. 432-437, DOI: 10.1109/IWFHR.2004.16, , 2004

    14. A Manchu Script Letters Dataset Creation Process with Simultaneous Letter Extraction and Labeling, A. D. Snowberger, C. H. Lee, accepted 2024, , 2024

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼