RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Adapting Language Models for Domain-Specific and Trustworthy Applications = 도메인 특화 및 신뢰 가능한 응용을 위한 언어 모델 적응 기법

    한글로보기

    https://www.riss.kr/link?id=T17535290

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Pre-trained language models, including large language models (LLMs), have become a general-purpose foundation for language understanding, reasoning, and generation. However, real-world use requires adaptation to specific users, domains, deployment constraints, and trustworthiness risks. This study investigates domain-specific and trustworthy language model adaptation through four methodological approaches. The first approach explores user-aware adaptation by using LLMs as proxies for student groups with different ability levels. The second approach extends toward practical deployment by automatically generating annotated training data from structured menu databases and using store-specific adapters for scalable voice-ordering systems. The third approach addresses reasoning over professional documents by adapting an LLM to construction standards and classifying sentence-level overlap, conflict, and neutrality with selective reasoning. The fourth approach examines trustworthiness by measuring and mitigating media outlet name bias in LLMs, showing that contextual source cues can systematically affect model behavior. Overall, the thesis argues that reliable language model adaptation requires user modeling, domain grounding, deployment-aware specialization, reasoning support, and bias-sensitive evaluation.
    번역하기

    Pre-trained language models, including large language models (LLMs), have become a general-purpose foundation for language understanding, reasoning, and generation. However, real-world use requires adaptation to specific users, domains, deployment con...

    Pre-trained language models, including large language models (LLMs), have become a general-purpose foundation for language understanding, reasoning, and generation. However, real-world use requires adaptation to specific users, domains, deployment constraints, and trustworthiness risks. This study investigates domain-specific and trustworthy language model adaptation through four methodological approaches. The first approach explores user-aware adaptation by using LLMs as proxies for student groups with different ability levels. The second approach extends toward practical deployment by automatically generating annotated training data from structured menu databases and using store-specific adapters for scalable voice-ordering systems. The third approach addresses reasoning over professional documents by adapting an LLM to construction standards and classifying sentence-level overlap, conflict, and neutrality with selective reasoning. The fourth approach examines trustworthiness by measuring and mitigating media outlet name bias in LLMs, showing that contextual source cues can systematically affect model behavior. Overall, the thesis argues that reliable language model adaptation requires user modeling, domain grounding, deployment-aware specialization, reasoning support, and bias-sensitive evaluation.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    사전학습된 언어 모델, 특히 초거대 언어 모델은 언어 이해, 추론, 생성 전반에 걸쳐 범용 기반 모델로 자리잡고 있다. 그러나 실제 환경에서의 활용을 위해서는 특정 사용자, 도메인, 배포 환경의 제약, 그리고 신뢰성 문제에 대한 적응이 필요하다. 본 연구는 도메인 특화 및 신뢰 가능한 언어 모델 적응을 네 가지 방법론적 접근을 통해 탐구한다. 첫 번째 접근에서는 서로 다른 능력 수준을 가진 학생 집단의 대리로 초거대 언어 모델을 활용함으로써 사용자 인지 적응 방법을 연구한다. 두 번째 접근에서는 구조화된 메뉴 데이터베이스로부터 라벨링된 학습 데이터를 자동 생성하고, 매장별 효율적 미세조정 어댑터를 활용하여 확장 가능한 음성 주문 시스템을 구축함으로써 실제 배포 환경으로 연구를 확장한다. 세 번째 접근에서는 건설 분야 기준 문서를 대상으로 초거대 언어 모델을 적응시켜 문장 수준의 중복, 상충, 무관 관계를 분류하고, 선택적 논리추론을 수행함으로써 전문 문서 추론 문제를 다룬다. 네 번째 접근에서는 초거대 언어 모델 내 언론사 이름 편향을 측정하고 완화함으로써 신뢰성 문제를 분석하며, 문맥적 출처 단서가 모델의 행동에 체계적인 영향을 미칠 수 있음을 보인다. 종합적으로 본 학위논문은 신뢰 가능한 언어 모델 적응을 위해 사용자 모델링, 도메인 적응, 배포 환경을 고려한 특화 아키텍처, 고도화된 동적 추론, 그리고 언어 모델의 편향을 고려한 평가가 필요함을 주장한다.
    번역하기

    사전학습된 언어 모델, 특히 초거대 언어 모델은 언어 이해, 추론, 생성 전반에 걸쳐 범용 기반 모델로 자리잡고 있다. 그러나 실제 환경에서의 활용을 위해서는 특정 사용자, 도메인, 배포 ...

    사전학습된 언어 모델, 특히 초거대 언어 모델은 언어 이해, 추론, 생성 전반에 걸쳐 범용 기반 모델로 자리잡고 있다. 그러나 실제 환경에서의 활용을 위해서는 특정 사용자, 도메인, 배포 환경의 제약, 그리고 신뢰성 문제에 대한 적응이 필요하다. 본 연구는 도메인 특화 및 신뢰 가능한 언어 모델 적응을 네 가지 방법론적 접근을 통해 탐구한다. 첫 번째 접근에서는 서로 다른 능력 수준을 가진 학생 집단의 대리로 초거대 언어 모델을 활용함으로써 사용자 인지 적응 방법을 연구한다. 두 번째 접근에서는 구조화된 메뉴 데이터베이스로부터 라벨링된 학습 데이터를 자동 생성하고, 매장별 효율적 미세조정 어댑터를 활용하여 확장 가능한 음성 주문 시스템을 구축함으로써 실제 배포 환경으로 연구를 확장한다. 세 번째 접근에서는 건설 분야 기준 문서를 대상으로 초거대 언어 모델을 적응시켜 문장 수준의 중복, 상충, 무관 관계를 분류하고, 선택적 논리추론을 수행함으로써 전문 문서 추론 문제를 다룬다. 네 번째 접근에서는 초거대 언어 모델 내 언론사 이름 편향을 측정하고 완화함으로써 신뢰성 문제를 분석하며, 문맥적 출처 단서가 모델의 행동에 체계적인 영향을 미칠 수 있음을 보인다. 종합적으로 본 학위논문은 신뢰 가능한 언어 모델 적응을 위해 사용자 모델링, 도메인 적응, 배포 환경을 고려한 특화 아키텍처, 고도화된 동적 추론, 그리고 언어 모델의 편향을 고려한 평가가 필요함을 주장한다.

    더보기

    목차 (Table of Contents)

    • Ⅰ Introduction 1
    • Ⅱ Large Language Models are Students at Various Levels: Zero-shot Question Difficulty Estimation 4
    • 2.1 Introduction 4
    • 2.2 Method 8
    • 2.2.1 Question-Solving with LLMs 9
    • Ⅰ Introduction 1
    • Ⅱ Large Language Models are Students at Various Levels: Zero-shot Question Difficulty Estimation 4
    • 2.1 Introduction 4
    • 2.2 Method 8
    • 2.2.1 Question-Solving with LLMs 9
    • 2.2.2 LLaSA 10
    • 2.2.2.1 LLM Clustering Module 10
    • 2.2.2.2 LLM Distribution Adjustment 12
    • 2.2.3 Zero-shot LLaSA 12
    • 2.3 Experiments 14
    • 2.3.1 Datasets 14
    • 2.3.2 Metrics 15
    • 2.3.3 Baselines 15
    • 2.3.4 Experimental Details 16
    • 2.3.5 QDE Results of LLaSA 16
    • 2.4 Analysis of LLaSA 20
    • 2.4.1 Question-Solving Based QDE of LLMs 20
    • 2.4.2 LLM Cluster Representation 22
    • 2.5 Related Works 24
    • 2.5.1 Question-Solving Skills of LLM 24
    • 2.5.2 Question Difficulty Estimation 24
    • 2.6 Conclusion 25
    • 2.7 Appendix 26
    • 2.7.1 Experimental Settings 26
    • 2.7.2 Implementation Details of LLaSA 31
    • 2.7.3 Additional Analysis 33
    • Ⅲ From Menus to the Interactive Food-Ordering Systems 40
    • 3.1 Introduction 40
    • 3.2 Related Works 44
    • 3.2.1 Conversational Interfaces for Task-Oriented Systems 44
    • 3.2.2 End-to-End Framework for Interface Development 45
    • 3.2.3 Dataset Generation and Augmentation Method 46
    • 3.3 Methodology 47
    • 3.3.1 Store-Specific Dataset Generation 48
    • 3.3.1.1 Enhancing Data Representation 48
    • 3.3.1.2 Template-Based Annotated Dataset Generation 49
    • 3.3.2 Store-Specific IC & SF Model Training 51
    • 3.3.2.1 Store-Specific Adapters 51
    • 3.3.2.2 Adapter Integration into a Unified Model 52
    • 3.3.3 Model Serving in Realistic Environments 52
    • 3.3.3.1 Recommendation Module for Unavailable Items 52
    • 3.3.3.2 End-to-End Adapter Update for New Stores 53
    • 3.4 Experiments 54
    • 3.4.1 Experimental Setup 54
    • 3.4.1.1 Store Information 54
    • 3.4.1.2 Test Datasets 56
    • 3.4.1.3 Data Generation Baselines 56
    • 3.4.1.4 Models 56
    • 3.4.1.5 Evaluation Metrics 57
    • 3.4.1.6 Implementation Details 57
    • 3.4.1.7 Details for Real-World Performance Evaluation 58
    • 3.4.2 Data Generation Details 59
    • 3.4.2.1 TUDA 59
    • 3.4.2.2 Ours 60
    • 3.4.2.3 Bllossom 60
    • 3.4.3 Experimental Results 61
    • 3.4.3.1 Annotated Dataset Generation 61
    • 3.4.3.2 Intent Classification and Slot Filling 63
    • 3.4.3.3 Effectiveness of Store-Specific Models 65
    • 3.4.3.4 Item Recommendation Module 66
    • 3.4.3.5 Real-World Performance Evaluation 68
    • 3.5 Conclusion and Discussion 68
    • 3.5.1 Limitations and Future Work 69
    • 3.5.2 Broader Impacts 70
    • Ⅳ Conflict and Overlap Classification in Construction Standards Using a Large Language Model 71
    • 4.1 Introduction 71
    • 4.2 Related Work 75
    • 4.3 Method 76
    • 4.3.1 Adapting LLM to Construction Domain 77
    • 4.3.2 Two-Step Classification of Overlap, Conflict, and Neutrality Using LLMs 78
    • 4.3.3 Interactive Interface for Construction Standard Analysis 81
    • 4.4 Experiment 82
    • 4.4.1 Dataset 82
    • 4.4.2 Evaluation Metrics 82
    • 4.4.3 Baselines 83
    • 4.4.4 Experimental Results 83
    • 4.5 Analysis 85
    • 4.6 Conclusion 86
    • 4.7 Appendix 87
    • 4.7.1 Paragraph-level Reasoning with Domain-adapted LLM 87
    • 4.7.2 Implementation Details of ITRM 89
    • 4.7.3 Prompt for COSLLM 89
    • 4.7.4 System Interfaces 90
    • 4.7.5 Implementation Details of COSLLM and Baselines 93
    • 4.7.6 Detailed Class-wise Performance 93
    • 4.7.7 CoT Reasoning Examples 94
    • Ⅴ Measuring and Mitigating Media Outlet Name Bias in Large Language Models 95
    • 5.1 Introduction 95
    • 5.2 Measuring Media Outlet Name Bias in LLMs 99
    • 5.2.1 Political Bias Prediction Shift 100
    • 5.2.2 The SIPS Metric 101
    • 5.2.3 Sentiment Shifts in Article Summarization 103
    • 5.3 Experimental Setup 104
    • 5.3.1 Representative Media Outlet Selection 104
    • 5.3.2 LLM Selection 104
    • 5.3.3 Political News Dataset Selection 105
    • 5.3.4 Implementation Details 105
    • 5.4 Results and Analysis 106
    • 5.4.1 News Article Political Bias Prediction 106
    • 5.4.2 News Article Summarization 112
    • 5.5 Mitigating Media Outlet Name Bias 114
    • 5.5.1 Prompt Optimization Strategies 114
    • 5.5.2 Results of the Prompt Optimization 115
    • 5.6 Related Works 116
    • 5.6.1 Political Bias in LLMs 116
    • 5.6.2 Applications of LLMs in the News Media and Political Science Domain 117
    • 5.7 Conclusion 118
    • 5.8 Appendix 119
    • 5.8.1 Comprehensive Ethical Considerations 119
    • 5.8.2 Interpretation of SIPS 120
    • 5.8.3 Implementation Details 122
    • 5.8.3.1 Entity-wise Summarization Analysis Method 122
    • 5.8.3.2 Detailed Criteria for LLM Selection 122
    • 5.8.3.3 Detailed LLM Prompts 123
    • 5.8.3.4 Detailed LLM Generation Configuration 123
    • 5.8.4 Additional Analysis Results for Article Bias Prediction 124
    • 5.8.4.1 Detailed Analysis by Media Bias Labels and Names 124
    • 5.8.4.2 Generated & Formulated Media Outlet Names 124
    • 5.8.4.3 Saliency Analysis Details 125
    • 5.8.5 Additional Analysis Results for Article Summarization 126
    • 5.8.5.1 Detailed Result of Entity-wise Summarization Analysis 126
    • 5.8.5.2 Detailed Result of Content-wise Summarization Analysis 126
    • 5.8.5.3 Human Evaluation Details 128
    • 5.8.5.4 Example of LLM Political Bias Perception 131
    • 5.8.6 Additional Prompt Optimization Results 132
    • 5.8.6.1 Detailed Prompt Optimization Results Across 6 Major Models 132
    • 5.8.6.2 Prompts Generated by Iterative Optimization 132
    • Ⅵ Concluding Remark 134
    • Ⅶ References 136
    • Ⅷ 국문 초록 165
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼