본 연구는 고도화된 한국어 법률 LLM 개발을 위한 로드맵 수립을 목표로, 한국 법률 도메인(성문법–판례 위계)에서 데이터 전략과 아키텍처 선택의 효과를 실증적으로 규명한다. 이를 위해 4...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
본 연구는 고도화된 한국어 법률 LLM 개발을 위한 로드맵 수립을 목표로, 한국 법률 도메인(성문법–판례 위계)에서 데이터 전략과 아키텍처 선택의 효과를 실증적으로 규명한다. 이를 위해 4...
본 연구는 고도화된 한국어 법률 LLM 개발을 위한 로드맵 수립을 목표로, 한국 법률 도메인(성문법–판례 위계)에서 데이터 전략과 아키텍처 선택의 효과를 실증적으로 규명한다. 이를 위해 4B 규모에서 Autoregressive(AR)와 Diffusion Language Model(DLM)을 동일 조건하에 통제 실험하였으며, 법령·판례·주석서의 계층적 Continual Pretraining과 변호사시험 선지를 분해한 Atomic Legal QA기반 SFT 파이프라인을 설계하였다. 실험 결과, 판례 중심 학습은 법률 문항 해결 능력을 높이나 일반 추론 성능을 저하시키는 Negative Transfer 현상을 야기함을 확인하였으며, 주석서 학습이 이를 상쇄하는 효과를 보였다. 특히, Atomic QA를 활용한 고도화된 SFT를 수행한 결과, DLM의 정확도는 33.48%까지 비약적으로 향상되어 일반 언어 모델 기준선 (27.67%)을 유의미하게 상회하였으며, 상용 모델(GPT-4o)과의 격차 또한 유의미하게 좁혔다. 이는 정교한 데이터 전략이 뒷받침될 때 기존 AR 위주의 생태계에서 Diffusion 모델 또한 법률 도메인에서 충분한 잠재력을 가질 수 있음을 실증한다. 이에 본 연구는 추론 중심 데이터와 RAG 등 고도화 방안을 제안한다. 결론적으로 본 연구는 강건한 한국어 법률 LLM 구축을 위해, 아키텍처의 내재적 특성보다 데이터의 질과 논리적 구조가 더 결정적인 요소임을 실증한다.
다국어 초록 (Multilingual Abstract)
This study aims to establish a roadmap for developing advanced Korean Legal Large Language Models (LLMs) by empirically investigating the effects of data strategies and architectural choices within the Korean legal domain. To this end, we conducted co...
This study aims to establish a roadmap for developing advanced Korean Legal Large Language Models (LLMs) by empirically investigating the effects of data strategies and architectural choices within the Korean legal domain. To this end, we conducted controlled experiments comparing Autoregressive (AR) models and Diffusion Language Models (DLM) at the 4B scale under identical conditions. We designed a training pipeline consisting of hierarchical Continual Pretraining on statutes, precedents, and commentaries, followed by Supervised Fine-Tuning (SFT) using ‘Atomic Legal QA,’ a dataset constructed by decomposing bar exam multiple-choice options.
The results confirm that precedent-focused training improves legal question- answering capabilities but induces ‘Negative Transfer,’ degrading general reasoning performance; however, training with commentaries was found to mitigate this trade-off. Notably, applying Atomic QA-based SFT significantly boosted the DLM’s accuracy on the Bar Exam to 33.48%, meaningfully surpassing the general language model baseline (27.67%) and narrowing the gap with commercial models. These results demonstrate that, even within an ecosystem dominated by AR models, Diffusion models possess sufficient potential for application in the legal domain when supported by sophisticated, reasoning-centric data strategies.
Accordingly, this study proposes advancement strategies such as reasoning- centric data construction and Retrieval-Augmented Generation (RAG). In conclusion, this research suggests that for building robust Korean Legal LLMs, the quality and logical structure of data are critical factors that can transcend architectural limitations.
목차 (Table of Contents)