RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    문서 기반 성별 예측을 위한 요인 추출 및 한글 문서에의 적용 연구 = A Study on Feature Extraction for Gender Prediction Using Text Documents and Their Application to Korean Corpus

    한글로보기

    https://www.riss.kr/link?id=A104124511

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    As gender information is required in diverse domains, gender prediction becomes an important research issue. Among gender pre-diction using various data types including image, video, and sensor data, gender prediction using text documents makes it possible to predict gender of users in social network or blog services using documents written by them. Gender prediction performance is closely related to the features extracted from documents and used for prediction. In this regard, we introduce feature extraction methods adopt-ed in previous gender prediction studies using text documents and investigate their application to a Korean corpus. We categorized the features into more than 40 types. Some of them, which can be applied to Korean corpus, were utilized for gender prediction using Ko-rean blog corpus. From the experiment, it can be concluded open dictionary features outperformed other lexical features and sematic feature is most effective for gender prediction.
    번역하기

    As gender information is required in diverse domains, gender prediction becomes an important research issue. Among gender pre-diction using various data types including image, video, and sensor data, gender prediction using text documents makes it pos...

    As gender information is required in diverse domains, gender prediction becomes an important research issue. Among gender pre-diction using various data types including image, video, and sensor data, gender prediction using text documents makes it possible to predict gender of users in social network or blog services using documents written by them. Gender prediction performance is closely related to the features extracted from documents and used for prediction. In this regard, we introduce feature extraction methods adopt-ed in previous gender prediction studies using text documents and investigate their application to a Korean corpus. We categorized the features into more than 40 types. Some of them, which can be applied to Korean corpus, were utilized for gender prediction using Ko-rean blog corpus. From the experiment, it can be concluded open dictionary features outperformed other lexical features and sematic feature is most effective for gender prediction.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 개인화된 추천 시스템과 같이 성별 정보를 필요로 하는 서비스가 증가함에 따라 사용자의 성별 예측은 주요 연구 주제로 각광받고 있다. 이미지, 동영상, 센서 등 다양한 데이터를 기반으로 성별 예측이 이루어지고 있으며, 이 중 SNS나 블로그의 글을 토대로 저자의 성별을 알아 낼 수 있다. 이때, 문서에서 추출된 요인의 종류에 따라 예측 성능이 달라진다고 알려져 있다. 따라서 본 연구에서는 기존 문서 기반 성별 예측 연구에서 사용된 요인의 종류 및 추출 방법론을 정리하고 이들의 한글 문서에의 적용 가능성을 살펴본다. 약 40종류 이상의 요인들이 정리되었으며, 이들 중 한글 문서에 적용 가능한 요인들을 선정하여 한글 블로그 문서에서 추출하였다. 이렇게 추출된 요인을 이용하여 성별 예측 실험을 수행하였으며 실험을 통해 열린 사전 요인과 의미 요인이 성별 예측에 유의미하다는 결론을 내릴 수 있었다.
    번역하기

    최근 개인화된 추천 시스템과 같이 성별 정보를 필요로 하는 서비스가 증가함에 따라 사용자의 성별 예측은 주요 연구 주제로 각광받고 있다. 이미지, 동영상, 센서 등 다양한 데이터를 기반...

    최근 개인화된 추천 시스템과 같이 성별 정보를 필요로 하는 서비스가 증가함에 따라 사용자의 성별 예측은 주요 연구 주제로 각광받고 있다. 이미지, 동영상, 센서 등 다양한 데이터를 기반으로 성별 예측이 이루어지고 있으며, 이 중 SNS나 블로그의 글을 토대로 저자의 성별을 알아 낼 수 있다. 이때, 문서에서 추출된 요인의 종류에 따라 예측 성능이 달라진다고 알려져 있다. 따라서 본 연구에서는 기존 문서 기반 성별 예측 연구에서 사용된 요인의 종류 및 추출 방법론을 정리하고 이들의 한글 문서에의 적용 가능성을 살펴본다. 약 40종류 이상의 요인들이 정리되었으며, 이들 중 한글 문서에 적용 가능한 요인들을 선정하여 한글 블로그 문서에서 추출하였다. 이렇게 추출된 요인을 이용하여 성별 예측 실험을 수행하였으며 실험을 통해 열린 사전 요인과 의미 요인이 성별 예측에 유의미하다는 결론을 내릴 수 있었다.

    더보기

    참고문헌 (Reference)

    1 김충일, "혼합필터링(Hybrid Filtering)을 활용한 교과목 추천시스템 개발" 엘지씨엔에스 14 (14): 71-82, 2015

    2 김소이, "특징적 단어 집합을 활용한 모바일 기기 내 성별 예측 프레임워크" 42 (42): 345-347, 2015

    3 류경민, "트위터 사용자의 성별, 나이, 위치 추론" 32 (32): 46-53, 2014

    4 정석봉, "인터넷 쇼핑몰의 고객분류 및 선호상품 추출을 위한 사회연결망분석 기반의 새로운 접근" 엘지씨엔에스 13 (13): 57-68, 2014

    5 김상채, "문체 분석을 활용한 한국어트위터 사용자의 연령대 및 성별 예측" 39 (39): 303-305, 2012

    6 송현제, "모빌리티와 소셜 미디어 텍스트에서 사용자 프로파일 식별" 한국정보과학회 19 (19): 393-397, 2013

    7 송현제, "다중 인스턴스 학습 기반 소셜 미디어 사용자 프로파일 식별" 한국정보과학회 40 (40): 233-240, 2013

    8 임명재, "나이브 베이지안을 사용한 성명에 대한 성별 구분 연구" 한국인터넷방송통신학회 13 (13): 155-159, 2013

    9 홍진성, "국가 현안 주제 선정을 위한 데이터 분석 기반 하이브리드 방법론" 엘지씨엔에스 13 (13): 97-111, 2014

    10 Nguyen, D., "Why Gender and Age Prediction from Tweets Is Hard: Lessons from a Crowdsourcing Experiment" 1950-1961, 2014

    1 김충일, "혼합필터링(Hybrid Filtering)을 활용한 교과목 추천시스템 개발" 엘지씨엔에스 14 (14): 71-82, 2015

    2 김소이, "특징적 단어 집합을 활용한 모바일 기기 내 성별 예측 프레임워크" 42 (42): 345-347, 2015

    3 류경민, "트위터 사용자의 성별, 나이, 위치 추론" 32 (32): 46-53, 2014

    4 정석봉, "인터넷 쇼핑몰의 고객분류 및 선호상품 추출을 위한 사회연결망분석 기반의 새로운 접근" 엘지씨엔에스 13 (13): 57-68, 2014

    5 김상채, "문체 분석을 활용한 한국어트위터 사용자의 연령대 및 성별 예측" 39 (39): 303-305, 2012

    6 송현제, "모빌리티와 소셜 미디어 텍스트에서 사용자 프로파일 식별" 한국정보과학회 19 (19): 393-397, 2013

    7 송현제, "다중 인스턴스 학습 기반 소셜 미디어 사용자 프로파일 식별" 한국정보과학회 40 (40): 233-240, 2013

    8 임명재, "나이브 베이지안을 사용한 성명에 대한 성별 구분 연구" 한국인터넷방송통신학회 13 (13): 155-159, 2013

    9 홍진성, "국가 현안 주제 선정을 위한 데이터 분석 기반 하이브리드 방법론" 엘지씨엔에스 13 (13): 97-111, 2014

    10 Nguyen, D., "Why Gender and Age Prediction from Tweets Is Hard: Lessons from a Crowdsourcing Experiment" 1950-1961, 2014

    11 Liu, W., "What’s in a Name? Using First Names as Features for Gender Inference in Twitter" 2013

    12 Tang, C., "What’s in a Name: A Study of Names, Gender Inference, and Gender Behavior in Facebook" 6637 : 344-356, 2011

    13 Heylighen, F., "Variation in the Contextuality of Language: An Empirical Measure" 7 (7): 293-340, 2002

    14 Mislove, A., "Understanding the Demographics of Twitter Users" 554-557, 2011

    15 Ikeda, K., "Twitter User Profiling Based on Text and Community Mining for Market Analysis" 51 : 35-47, 2013

    16 Goswami, S., "Stylometric Analysis of Bloggers’ Age and Gender" 214-217, 2009

    17 Quinlan, J. R., "Simplifying Decision Trees" 51 (51): 497-510, 1999

    18 Peersman, C., "Predicting Age and Gender in Online Social Networks" 37-44, 2011

    19 Schwartz, H. A., "Personality, Gender, and Age in the Language of Social Media: The Open-Vocabulary Approach" 8 (8): 2013

    20 Garera, N., "Modeling Latent Biographic Attributes in Conversational Genres" 2 : 2009

    21 Argamon, S., "Mining the Blogosphere: Age, Gender and the Varieties of Self-Expression" 12 (12): 2007

    22 Dumais, S. T., "Latent Semantic Analysis" 38 (38): 188-230, 2004

    23 Blei, D. M., "Latent Dirichlet Allocation" 3 : 993-1022, 2003

    24 Lakoff, R., "Language and Woman’s Place" 2 (2): 45-80, 1973

    25 Bi, B., "Inferring the Demographics of Search Users" 131-140, 2013

    26 Otterbacher, J., "Inferring Gender of Movie Reviewers: Exploiting Writing Style, Content and Metadata" 369-378, 2010

    27 Mukherjee, A., "Improving Gender Classification of Blog Authors" 207-217, 2010

    28 Huang, F., "Identifying Gender of Microblog Users Based on Message Mining" 8485 : 488-493, 2014

    29 Argamon, S., "Gender, Genre, and Writing Style in Formal Written Texts" 23 (23): 321-346, 2003

    30 Ciot, M., "Gender Inference of Twitter Users in Non-English Contexts" 1136-1145, 2013

    31 Bamman, D., "Gender Identity and Lexical Variation in Social Media" 18 (18): 135-160, 2014

    32 Cheng, N., "Gender Identification from E-Mails" 154-158, 2009

    33 Igarashi, T., "Gender Differences in Social Network Development via Mobile Phone Text Messages: A Longitudinal Study" 22 (22): 691-713, 2005

    34 Shahana, P. H., "Feature Selection Techniques for Gender Prediction from Blogs" 355-359, 2014

    35 Weren, E. R., "Examining Multiple Features for Author Profiling" 5 (5): 266-279, 2014

    36 Alowibdi, J. S., "Empirical Evaluation of Profile Characteristics for Gender Classification on Twitter" 1 : 365-369, 2013

    37 Schler, J., "Effects of Age and Gender on Blogging" 6 : 199-205, 2006

    38 Burger, J. D., "Discriminating Gender on Twitter" 1301-1309, 2011

    39 Li, L., "Discriminating Gender on Chinese Microblog: A Study of Online Behaviour, Writing Style and Preferred Vocabulary" 812-817, 2014

    40 Hu, J., "Demographic Prediction Based on User’s Browsing Behavior" 151-160, 2007

    41 Rao, D., "Classifying Latent User Attributes in Twitter" 37-44, 2010

    42 Kucukyilmaz, T., "Chat Mining for Gender Prediction" 4243 : 274-283, 2006

    43 Bergsma, S., "Broadly Improving User Classification via Communication-Based Name and Location Clustering on Twitter" 1010-1019, 2013

    44 최예림, "BCI 기반 무인기 조종사 집중도 유지 지상 통제 프레임워크" 엘지씨엔에스 12 (12): 101-115, 2013

    45 Argamon, S., "Automatically Profiling the Author of an Anonymous Text" 52 (52): 119-123, 2009

    46 Mikros, G. K., "Authorship Attribution and Gender Identification in Greek Blogs" 21-32, 2012

    47 Estival, D., "Author Profiling for English and Arabic Emails" 1 (1): 1-22, 2008

    48 Deitrick, W., "Author Gender Prediction in an Email Stream Using Neural Networks" 4 (4): 169-175, 2012

    49 Cheng, N., "Author Gender Identification from Text" 8 (8): 78-88, 2011

    50 Pham, D. D., "Author Profiling for Vietnamese Blogs" 190-194, 2009

    51 Estival, D., "Author Profiling for English Emails" 263-272, 2007

    52 Alsmearat, K., "An Extensive Study of the Bag-of-Words Approach for Gender Identification of Arabic Articles" 601-608, 2014

    53 Hayashi, J., "Age and Gender Estimation from Facial Image Processing" 1 : 2002

    54 Villegas, M. P., "A Spanish Text Corpus for the Author Profiling Task" 2014

    55 Flesch, R., "A New Readability Yardstick" 32 (32): 221-233, 1948

    56 Dale, E., "A Formula for Predicting Readability: Instructions" 27 (27): 37-54, 1948

    57 Coleman, M., "A Computer Readability Formula Designed for Machine Scoring" 60 (60): 283-284, 1975

    58 Yu, S., "A Study on Gait-Based Gender Classification" 18 (18): 1905-1910, 2009

    59 Zheng, R., "A Framework for Authorship Identification of Online Messages: Writing-Style Features and Classification Techniques" 57 (57): 378-393, 2006

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    인용정보 인용지수 설명보기

    학술지 이력

    학술지 이력
    연월일 이력구분 이력상세 등재구분
    2020 평가 신규평가 신청대상 (신규평가)
    2019-12-01 등재 등재후보 탈락 (계속평가)
    2018-12-01 등재 등재후보로 하락 (계속평가) KCI등재후보
    2015-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2011-01-01 등재 등재학술지 유지 (등재유지) KCI등재
    2008-01-01 등재 등재학술지 선정 (등재후보2차) KCI등재
    2007-01-01 등재 등재후보 1차 PASS (등재후보1차) KCI등재후보
    2005-01-01 등재 등재후보학술지 선정 (신규평가) KCI등재후보
    더보기

    학술지 인용정보

    학술지 인용정보
    기준연도 WOS-KCI 통합IF(2년) KCIF(2년) KCIF(3년)
    2016 0.8 0.8 0.73
    KCIF(4년) KCIF(5년) 중심성지수(3년) 즉시성지수
    0.79 0.86 0.972 0.06
    더보기

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼