본 연구는 CommonCrawl WET 데이터로부터 언어모델 훈련을 위한 대규모 한국어 웹 텍스트 데이터 정제 기법 중 라인 필터링 기법을 개선하여 사전학습 언어모델(Pretrained Language Model)의 단답형 문...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
본 연구는 CommonCrawl WET 데이터로부터 언어모델 훈련을 위한 대규모 한국어 웹 텍스트 데이터 정제 기법 중 라인 필터링 기법을 개선하여 사전학습 언어모델(Pretrained Language Model)의 단답형 문...
본 연구는 CommonCrawl WET 데이터로부터 언어모델 훈련을 위한 대규모 한국어 웹 텍스트 데이터 정제 기법 중 라인 필터링 기법을 개선하여 사전학습 언어모델(Pretrained Language Model)의 단답형 문항 문제풀이 성능을 개선한다. CommonCrawl은 2007년부터 수집한 2500억 개의 웹페이지를 제공하여 언어모델 학습을 위한 방대한 텍스트 자원으로 활용된다. 특히 WET 형식은 HTML이 아닌 텍스트만 포함하고 있어, 원본 HTML을 직접 처리하기 어려운 학계 및 비영리 기관에서도 널리 활용되고 있다. 그러나 WET 데이터에는 모델 훈련에 불필요한 텍스트 라인이 많아, 언어모델 성능을 개선하는 고품질 학습 데이터를 얻기 위해서는 효과적인 라인 필터링 기법이 필수적이다. 널리 사용되는 문장부호 기반 라인 필터링 기법이 있으나, 본 연구는 이 기법이 언어모델의 단답형 문제 풀이 성능 저하를 유발함을 실험적으로 보인다.
본 연구는 CommonCrawl WET 형식으로 주어진 문서에서, 각 라인의 특성 및 특성이 문서에 분포한 패턴을 활용하여 라인 필터링을 수행하는 기법을 제안한다. 먼저, 본 연구는 문장부호로 끝나지 않는 라인을 제거하는 필터링 규칙이 사전학습된 언어모델의 단답형 문항 문제풀이 성능 하락을 유발함을 지적하고, 이를 개선하는 문장부호 패턴 기반 라인 필터링 기법을 제안한다. 본 연구에서 제안하는 패턴 기반의 라인 필터링 기법을 적용하여 훈련한 모델은 CCNet, Dolma의 기법을 적용하여 훈련한 모델에 비해 KorQuAD(v1) 점수가 각각 +2.3점, +26.1점 상승하며, 벤치마크 점수 평균이 각각 +2.0점, +6.1점 상승한다.
다국어 초록 (Multilingual Abstract)
This study improves the line filtering rule in the preprocessing pipeline for constructing large-scale Korean web text datasets from CommonCrawl WET data, enhancing the generative question answering performance of language models. CommonCrawl provides...
This study improves the line filtering rule in the preprocessing pipeline for constructing large-scale Korean web text datasets from CommonCrawl WET data, enhancing the generative question answering performance of language models. CommonCrawl provides a vast text resource of over 250 billion web pages collected since 2007, widely used for language model training. While the WET format that only contains text is accessible to research institutions, the data contains many lines of text that are irrelevant for model training, making effective line filtering essential for obtaining high-quality training data. Although punctuation-based line filtering is widely used, this study empirically demonstrates that such methods can degrade performance on generative tasks.
We propose a new line filtering method that leverages both the characteristics of individual lines and the distributional patterns of those characteristics within each document. We first identify that Punctuation Mark Rule adversely affects generative task performance and then propose a novel rule to address this issue. A model trained using our proposed pattern-based filtering method achieves +2.3 and +26.1 point gains on the KorQuAD (v1) benchmark, +2.0 and +6.1 points on average, compared to models trained with the CCNet and Dolma methods, respectively.
목차 (Table of Contents)