RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    단어 쓰임새 정보와 신경망을 활용한 한국어 Hedge 인식 = Korean Hedge Detection Using Word Usage Information and Neural Networks

    한글로보기

    https://www.riss.kr/link?id=A104306380

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    In this paper, we try to classify Korean hedge sentences, which are regarded as not important since they express uncertainties or personal assumptions. Through previous researches to English language, we found dependency information of words has been one of important features in hedge classification, but not used in Korean researches. Additionally, we found that word embedding vectors include the word usage information. We assume that the word usage information could somehow represent the dependency information. Therefore, we utilized word embedding and neural networks in hedge sentence classification. We used more than one and half million sentences as word embedding dataset and also manually constructed 12,517-sentence hedge classification dataset obtained from online news. We used SVM and CRF as our baseline systems and the proposed system outperformed SVM by 7.2%p and also CRF by 1.2%p. This indicates that word usage information has positive impacts on Korean hedge classification.
    번역하기

    In this paper, we try to classify Korean hedge sentences, which are regarded as not important since they express uncertainties or personal assumptions. Through previous researches to English language, we found dependency information of words has been ...

    In this paper, we try to classify Korean hedge sentences, which are regarded as not important since they express uncertainties or personal assumptions. Through previous researches to English language, we found dependency information of words has been one of important features in hedge classification, but not used in Korean researches. Additionally, we found that word embedding vectors include the word usage information. We assume that the word usage information could somehow represent the dependency information. Therefore, we utilized word embedding and neural networks in hedge sentence classification. We used more than one and half million sentences as word embedding dataset and also manually constructed 12,517-sentence hedge classification dataset obtained from online news. We used SVM and CRF as our baseline systems and the proposed system outperformed SVM by 7.2%p and also CRF by 1.2%p. This indicates that word usage information has positive impacts on Korean hedge classification.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    본 논문에서는 한국어 문장을 대상으로 불확실한 사실이나 개인적인 추측으로 인해 중요하지 않다고 판단되는 문장, 즉 Hedge 문장들을 분류해 내고자 한다. 기존 영어권 연구에서는 Hedge 문장들을 분류할 때 단어의 의존관계 정보가 여러 형태로 활용되고 있으나, 한국어 연구에서는 사용되고 있지 않음을 확인하였다. 또 기존의 워드 임베딩(Word Embedding) 기법에서 단어의 쓰임새 정보가 학습된다는 점을 인지하였다. 단어의 쓰임새 정보가 어느 정도 의존관계를 표현할 수 있을 것으로 보고 워드 임베딩 정보를 Hedge 분류 실험에 적용하였다. 기존에 많이 사용되던 SVM과 CRF를 baseline 시스템으로 활용하였고 워드 임베딩과 신경망을 사용하여 비교실험을 하였다. 워드임베딩 데이터는 세종데이터와 온라인에서 수집된 데이터를 합하여 총 150여만 문장을 사용하였고 Hedge 분류 데이터는 수작업으로 구축한 12,517 문장의 뉴스데이터를 사용하였다. 워드 임베딩을 사용한 시스템이 SVM보다 7.2%p, CRF보다 1.6%p 좋은 성능을 내는 것을 확인하였다. 이는 단어의 쓰임새 정보가 한국어 Hedge 분류에서 긍정적인 영향을 미친다는 것을 의미한다.
    번역하기

    본 논문에서는 한국어 문장을 대상으로 불확실한 사실이나 개인적인 추측으로 인해 중요하지 않다고 판단되는 문장, 즉 Hedge 문장들을 분류해 내고자 한다. 기존 영어권 연구에서는 Hedge 문...

    본 논문에서는 한국어 문장을 대상으로 불확실한 사실이나 개인적인 추측으로 인해 중요하지 않다고 판단되는 문장, 즉 Hedge 문장들을 분류해 내고자 한다. 기존 영어권 연구에서는 Hedge 문장들을 분류할 때 단어의 의존관계 정보가 여러 형태로 활용되고 있으나, 한국어 연구에서는 사용되고 있지 않음을 확인하였다. 또 기존의 워드 임베딩(Word Embedding) 기법에서 단어의 쓰임새 정보가 학습된다는 점을 인지하였다. 단어의 쓰임새 정보가 어느 정도 의존관계를 표현할 수 있을 것으로 보고 워드 임베딩 정보를 Hedge 분류 실험에 적용하였다. 기존에 많이 사용되던 SVM과 CRF를 baseline 시스템으로 활용하였고 워드 임베딩과 신경망을 사용하여 비교실험을 하였다. 워드임베딩 데이터는 세종데이터와 온라인에서 수집된 데이터를 합하여 총 150여만 문장을 사용하였고 Hedge 분류 데이터는 수작업으로 구축한 12,517 문장의 뉴스데이터를 사용하였다. 워드 임베딩을 사용한 시스템이 SVM보다 7.2%p, CRF보다 1.6%p 좋은 성능을 내는 것을 확인하였다. 이는 단어의 쓰임새 정보가 한국어 Hedge 분류에서 긍정적인 영향을 미친다는 것을 의미한다.

    더보기

    참고문헌 (Reference)

    1 정주석, "한국어 Hedge 문장 인식을 위한 태깅 말뭉치 및 단서어구 패턴 구축" 한국지능시스템학회 21 (21): 761-766, 2011

    2 임미영, "의미 정보가 강화된 워드 임베딩을 통한 감성 분석" 사단법인 인문사회과학기술융합학회 7 (7): 321-329, 2017

    3 O. Täckström, "Uncertainty detection as approximate max-margin sequence labelling" 84-91, 2010

    4 "TensorFlow"

    5 R. Rehurek, "Software Framework for Topic Modelling With Large Corpora" 46-50, 2010

    6 E. Velldal, "Resolving speculation: MaxEnt cue classification and dependency-based scope rules" 48-55, 2010

    7 R. Morante, "Memory-based resolution of in-sentence scopes of hedge cues" 40-47, 2010

    8 D. Tang, "Learning Sentiment-Specific Word Embedding for Twitter Sentiment Classification" 1 : 1555-1565, 2014

    9 E. R. Fernandes, "Hedge detection using the RelHunter approach" 64-69, 2010

    10 S. Zhang, "Hedge detection and scope finding by sequence labeling with normalized feature selection" 92-99, 2010

    1 정주석, "한국어 Hedge 문장 인식을 위한 태깅 말뭉치 및 단서어구 패턴 구축" 한국지능시스템학회 21 (21): 761-766, 2011

    2 임미영, "의미 정보가 강화된 워드 임베딩을 통한 감성 분석" 사단법인 인문사회과학기술융합학회 7 (7): 321-329, 2017

    3 O. Täckström, "Uncertainty detection as approximate max-margin sequence labelling" 84-91, 2010

    4 "TensorFlow"

    5 R. Rehurek, "Software Framework for Topic Modelling With Large Corpora" 46-50, 2010

    6 E. Velldal, "Resolving speculation: MaxEnt cue classification and dependency-based scope rules" 48-55, 2010

    7 R. Morante, "Memory-based resolution of in-sentence scopes of hedge cues" 40-47, 2010

    8 D. Tang, "Learning Sentiment-Specific Word Embedding for Twitter Sentiment Classification" 1 : 1555-1565, 2014

    9 E. R. Fernandes, "Hedge detection using the RelHunter approach" 64-69, 2010

    10 S. Zhang, "Hedge detection and scope finding by sequence labeling with normalized feature selection" 92-99, 2010

    11 X. Li, "Exploiting rich features for detecting hedges and their scope" 78-83, 2010

    12 H. Zhou, "Exploiting multi-features to detect hedges and their scope in biomedical texts" 106-113, 2010

    13 D. Lewis, "Evaluating Text Categorization" 91 : 312-318, 1991

    14 T. Milokov, "Distributed Representations of Words and Phrases and Their Compositionality" 3111-3119, 2013

    15 A. Vlachos, "Detecting speculative language using syntactic dependencies and logistic regression" 18-25, 2010

    16 F. Ji, "Detecting hedge cues and their scopes with average perceptron" 32-39, 2010

    17 M. Rei, "Combining manual rules and supervised learning for hedge cue and scope detection" 56-63, 2010

    18 B. Tang, "A cascade method for detecting hedges and their scope in natural language text" 13-17, 2010

    19 S. Kang, "A Comparison of Classifiers for Detecting Hedges" 251-257, 2011

    더보기

    동일학술지(권/호) 다른 논문

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    인용정보 인용지수 설명보기

    학술지 이력

    학술지 이력
    연월일 이력구분 이력상세 등재구분
    2020 평가 신규평가 신청대상 (신규평가)
    2019-12-01 등재 등재 탈락 (기타)
    2019-01-01 등재 등재학술지 유지 (계속평가) KCI등재
    2016-01-01 등재 등재학술지 선정 (계속평가) KCI등재
    2014-01-01 등재 등재후보학술지 선정 (신규평가) KCI등재후보
    더보기

    학술지 인용정보

    학술지 인용정보
    기준연도 WOS-KCI 통합IF(2년) KCIF(2년) KCIF(3년)
    2016 0.33 0.33 0.32
    KCIF(4년) KCIF(5년) 중심성지수(3년) 즉시성지수
    0.33 0.32 0.407 0.14
    더보기

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼