RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    Ternary Decomposition and Dictionary Extension for Khmer Word Segmentation

    한글로보기

    https://www.riss.kr/link?id=A101993293

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    In this paper, we proposed a dictionary extension and a ternary decomposition technique to improve the effectiveness of Khmer word segmentation. Most word segmentation approaches depend on a dictionary. However, the dictionary being used is not fully reliable and cannot cover all the words of the Khmer language. This causes an issue of unknown words or out-of-vocabulary words. Our approach is to extend the original dictionary to be more reliable with new words. In addition, we use ternary decomposition for the segmentation process. In this research, we also introduced the invisible space of the Khmer Unicode (char\u200B) in order to segment our training corpus. With our segmentation algorithm, based on ternary decomposition and invisible space, we can extract new words from our training text and then input the new words into the dictionary. We used an extended wordlist and a segmentation algorithm regardless of the invisible space to test an unannotated text. Our results remarkably outperformed other approaches. We have achieved 88.8%, 91.8% and 90.6% rates of precision, recall and F-measurement.
    번역하기

    In this paper, we proposed a dictionary extension and a ternary decomposition technique to improve the effectiveness of Khmer word segmentation. Most word segmentation approaches depend on a dictionary. However, the dictionary being used is not fully ...

    In this paper, we proposed a dictionary extension and a ternary decomposition technique to improve the effectiveness of Khmer word segmentation. Most word segmentation approaches depend on a dictionary. However, the dictionary being used is not fully reliable and cannot cover all the words of the Khmer language. This causes an issue of unknown words or out-of-vocabulary words. Our approach is to extend the original dictionary to be more reliable with new words. In addition, we use ternary decomposition for the segmentation process. In this research, we also introduced the invisible space of the Khmer Unicode (char\u200B) in order to segment our training corpus. With our segmentation algorithm, based on ternary decomposition and invisible space, we can extract new words from our training text and then input the new words into the dictionary. We used an extended wordlist and a segmentation algorithm regardless of the invisible space to test an unannotated text. Our results remarkably outperformed other approaches. We have achieved 88.8%, 91.8% and 90.6% rates of precision, recall and F-measurement.

    더보기

    목차 (Table of Contents)

    • Abstract
    • 1. Introduction
    • 2. Khmer Language Overview
    • 3. Research Reviews
    • 4. Proposed Approach
    • Abstract
    • 1. Introduction
    • 2. Khmer Language Overview
    • 3. Research Reviews
    • 4. Proposed Approach
    • 5. Experiment
    • 6. Conclusion
    • References
    더보기

    참고문헌 (Reference)

    1 Chea, S., "Word Bigram Vs Orthographic Syllable Bigram in Khmer Word Segmentation"

    2 Seng, S., "Which Units for acoustic and language modelling for Khmer automatic speech recognition?"

    3 Van, C., "Query Expansion for Khmer Information Retrieval" 80-87, 2010

    4 Seng, S., "Multiple Text Segmentation for Statistical Language Modelling" France 2MICA Center 2009

    5 Mohri, M. F., "Lecture Notes in Computer Science" Springer 144-158, 1998

    6 Channa, V., "Khmer Word Segmentation and Out -of-Vocabulary Words Detection Using Collection Measurement of Repeated Characters Subsequences"

    7 Khin, S., "Khmer Grammar" Royal Academy of Cambodia 2007

    8 "Khmer Dictionary" Royal Academy of Cambodia 2005

    9 Nevill-Manning, C. G., "Identifying Hierarchical Structure in Sequences A linear-time algorithm" 7 (7): 67-82, 1997

    10 Nou, C., "Hybrid Approach for Khmer Unknown Word POS Guessing"

    1 Chea, S., "Word Bigram Vs Orthographic Syllable Bigram in Khmer Word Segmentation"

    2 Seng, S., "Which Units for acoustic and language modelling for Khmer automatic speech recognition?"

    3 Van, C., "Query Expansion for Khmer Information Retrieval" 80-87, 2010

    4 Seng, S., "Multiple Text Segmentation for Statistical Language Modelling" France 2MICA Center 2009

    5 Mohri, M. F., "Lecture Notes in Computer Science" Springer 144-158, 1998

    6 Channa, V., "Khmer Word Segmentation and Out -of-Vocabulary Words Detection Using Collection Measurement of Repeated Characters Subsequences"

    7 Khin, S., "Khmer Grammar" Royal Academy of Cambodia 2007

    8 "Khmer Dictionary" Royal Academy of Cambodia 2005

    9 Nevill-Manning, C. G., "Identifying Hierarchical Structure in Sequences A linear-time algorithm" 7 (7): 67-82, 1997

    10 Nou, C., "Hybrid Approach for Khmer Unknown Word POS Guessing"

    11 Puthick, H., "Development of a Khmer Spell Checker Based on a Hidden Markov Model" Australian National University 2005

    12 Chea, S., "Detection and Correction of Homophonous Error Word for Khmer Language"

    13 Thanopoulos, A., "Comparative Evaluation of Collocation Extraction Metrics"

    14 Huffman, F. E., "Cambodian System of Writing and beginning reader with Drills and Glossary" Yale University Press 1970

    15 Seng, S., "Boosting N-gram Coverage for Unsegmented Languages Using Multiple Text Segmentation Approach" 1-7, 2010

    16 Church, K. W., "A Status Report on ACL/DCL" 84-91, 1991

    17 Shannon, E., "A Mathematical Theory of Communication" 27 : 379-423, 1948

    더보기

    동일학술지(권/호) 다른 논문

    동일학술지 더보기

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    인용정보 인용지수 설명보기

    학술지 이력

    학술지 이력
    연월일 이력구분 이력상세 등재구분
    2026 평가 재인증평가 신청대상 (재인증)
    2020-04-01 학회명변경 한글명 : 한국데이타베이스학회 -> 한국데이터전략학회
    영문명 : 미등록 -> Korea Data Strategy Society
    KCI등재
    2020-01-01 등재 등재학술지 유지 (재인증) KCI등재
    2017-01-01 등재 등재학술지 유지 (계속평가) KCI등재
    2013-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2010-06-22 학술지명변경 한글명 : Journal of Information Technology Applications & Menagement -> Journal of Information Technology Applications & Management
    외국어명 : Journal of Information Technology Applications & Menagement -> Journal of Information Technology Applications & Management
    KCI등재
    2010-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2008-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2005-01-01 등재 등재학술지 선정 (등재후보2차) KCI등재
    2004-01-01 등재 등재후보 1차 PASS (등재후보1차) KCI등재후보
    2002-01-01 등재 등재후보학술지 선정 (신규평가) KCI등재후보
    더보기

    학술지 인용정보

    학술지 인용정보
    기준연도 WOS-KCI 통합IF(2년) KCIF(2년) KCIF(3년)
    2016 0.39 0.39 0.48
    KCIF(4년) KCIF(5년) 중심성지수(3년) 즉시성지수
    0.59 0.56 0.673 0.18
    더보기

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼