본 연구는 제2언어 학습자의 중국어 학습에서 난도가 높은 문법 범주 중 하나인 결과보어 구문의 자동 추출 문제를 해결하기 위해, 대규모 언어 모델(LLM)에 기반한 결과보어 특화 데이터셋 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=A110186302
2026
Korean
KCI등재
학술저널
383-414(32쪽)
0
상세조회0
다운로드본 연구는 제2언어 학습자의 중국어 학습에서 난도가 높은 문법 범주 중 하나인 결과보어 구문의 자동 추출 문제를 해결하기 위해, 대규모 언어 모델(LLM)에 기반한 결과보어 특화 데이터셋 ...
본 연구는 제2언어 학습자의 중국어 학습에서 난도가 높은 문법 범주 중 하나인 결과보어 구문의 자동 추출 문제를 해결하기 위해, 대규모 언어 모델(LLM)에 기반한 결과보어 특화 데이터셋 구축 프레임워크를 제안한다. 중국어 결과보어는 동사와 보어가 결합해 복합동사처럼 기능한다. 또한 결과보어 구문은 ‘동사+동사' 또는 ‘동사+형용사' 형태로 구성되지만, 동일한 표면 구조가 항상 결과보어를 의미하는 것은 아니기에 표면 형식 기반 탐지 방식으로는 신뢰성 있는 구별이 어렵다. 범용 중국어 형태소 분석기(HanLP, Jieba, Stanza 등), CCL, BCC와 같은 코퍼스 검색 역시 결과보어 구분의 체계적 추출에 한계를 지니고 있어 기존 방식만으로는 신뢰도 높은 자동 추출이 어렵다. 이에 본 연구는 결과보어의 언어학적 특징을 반영한 판정 기준을 바탕으로 결과보어 구조를 자동 식별할 수 있는 모델을 설계하고자 한다. 이러한 프레임워크는 대규모 코퍼스에서 해당 구문을 정확하게 추출하여 미검출 및 오검출을 최소화하는 데 중점을 둔다. 이러한 접근은 고빈도 유형 중심이었던 기존 연구의 한계를 넘어, 다양한 결과보어를 포괄하는 대규모 데이터셋 구축을 가능하게 한다. 이는 중국어 문법 연구 및 제2언어 학습자 대상 교육 자료 개발에도 실질적 기여를 할 것으로 기대된다.
다국어 초록 (Multilingual Abstract)
This study proposes a large language model (LLM)–based framework for constructing a specialized dataset of Chinese resultative complement constructions (RCCs) to overcome the limitations of existing automatic extraction methods. Despite being a crit...
This study proposes a large language model (LLM)–based framework for constructing a specialized dataset of Chinese resultative complement constructions (RCCs) to overcome the limitations of existing automatic extraction methods. Despite being a critical yet challenging category for second language learners, RCCs are difficult to identify accurately because their "verb + verb/adjective" surface structures often overlap with non-resultative compounds, leading to significant structural ambiguity. While conventional morphological analyzers like HanLP and Jieba, as well as corpora such as CCL and BCC, often struggle with high-precision extraction, this research introduces a model that utilizes linguistically informed diagnostic criteria to capture the defining properties of RCCs. By prioritizing precision and minimizing error rates across large-scale corpora, this framework moves beyond high-frequency pattern matching to enable the construction of a comprehensive, diverse RCC dataset. Ultimately, this work contributes a robust computational approach to Chinese linguistic research and provides a foundational resource for developing advanced pedagogical materials.
杜甫 詩의 突兀·翻騰·逸宕之筆과 章法의 융합 고찰 ― 渾漫與와 詩律細 논의를 덧붙여
曾國藩 門下 弟子들의 文學 交流와 桐城 學術의 發展— 張裕釗, 吳汝綸, 黎庶昌을 중심으로
‘모성’에 갇힌 현대중국의 자화상― 모옌 《풍유비둔(豊乳肥臀)》을 읽는 하나의 방법