RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Enhancing Language Model Reliability Through Task-Specific Abstention Mechanisms = 상황별 응답 거부를 통한 언어모델 신뢰성 향상 연구

    한글로보기

    https://www.riss.kr/link?id=T17452093

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    언어모델은 자연어 처리 분야에서 눈부신 발전을 이루었으나, 부정확하거나 사용자 의도에 부합하지 않는 응답을 빈번히 생성하는 경향으로 인해 여전히 신뢰성 측면에서 근본적인 한계를 지닌다.
    본 논문은 이러한 문제를 해결하기 위해, 모델이 신뢰하기 어렵거나 의도에서 벗어난 응답을 명시적으로 거부(abstain)하는 관점에서 접근한다.
    특히, 응답의 신뢰성에 대한 조건과 요구사항은 주어진 상황에 따라 크게 상이함으로, 본 연구는 주어진 시나리오의 특성에 적합하게 설계된 상황별 응답 거부(task-specific abstention) 접근법을 탐구한다.

    이에 따라, 본 논문은 세 가지 대표적인 시나리오를 중심으로 상황별 응답 거부 접근법을 제안한다.
    첫째, Universal Domain Adaptation (UniDA)은 미지의 외부 도메인에서 유입되는 입력에 대해 적응(adaptation)할지 혹은 응답을 거부할지를 판단해야 하는 문제다.
    본 연구는 자연어처리 분야 최초의 UniDA 통합 평가 환경을 구축하고, 기존 연구들이 해당 환경에서 어떻게 동작하는지 체계적으로 분석하였다.
    둘째로, 사용자로부터 모호한(ambiguous) 질의가 제공되었을 때를 탐구한다.
    모호한 질의는 하나의 질문이 복수의 유효한 해석을 가질 수 있기 때문에, 사용자 의도에 벗어날 수 있는 임의의 답변을 제공하는 것 보다 사용자로부터 명확한 의도를 확인하는 것이 중요하다.
    본 연구는 질의 내 모호성을 탐지하고 명확화(clarification) 요청을 생성함으로써 사용자의 의도에 벗어나는 딥변을 방지하는 학습 프레임워크를 제안한다.
    마지막으로, 외부 지식을 활용한 질의응답 시나리오를 다룬다.
    이런 상황에서 모델은 입력으로 제공된 외부 지식과 사전학습된 모델의 지식을 모두 사용할 수 있다.
    사용자로부터 주어진 질의에 답변하기 위한 명확한 근거가 부재한 경우 응답을 거부할 수 있도록, 외부 지식과 사전학습된 지식에 동적으로 가중치는 부여하는 디코딩 방법론을 제안한다.

    다양한 실험을 통해, 본 논문은 상황에 따른 답변 거부 메커니즘이 언어모델의 신뢰성, 강건성 및 안전성을 크게 향상시킴을 보였다.
    종합하면, 본 연구는 언어모델의 응답 거부 능력 향상을 위한 포괄적 프레임워크를 제시하며, 실세계 환경에서 신뢰 가능하고 책임 있는 언어모델의 활용 기반을 마련한다.
    번역하기

    언어모델은 자연어 처리 분야에서 눈부신 발전을 이루었으나, 부정확하거나 사용자 의도에 부합하지 않는 응답을 빈번히 생성하는 경향으로 인해 여전히 신뢰성 측면에서 근본적인 한계를 ...

    언어모델은 자연어 처리 분야에서 눈부신 발전을 이루었으나, 부정확하거나 사용자 의도에 부합하지 않는 응답을 빈번히 생성하는 경향으로 인해 여전히 신뢰성 측면에서 근본적인 한계를 지닌다.
    본 논문은 이러한 문제를 해결하기 위해, 모델이 신뢰하기 어렵거나 의도에서 벗어난 응답을 명시적으로 거부(abstain)하는 관점에서 접근한다.
    특히, 응답의 신뢰성에 대한 조건과 요구사항은 주어진 상황에 따라 크게 상이함으로, 본 연구는 주어진 시나리오의 특성에 적합하게 설계된 상황별 응답 거부(task-specific abstention) 접근법을 탐구한다.

    이에 따라, 본 논문은 세 가지 대표적인 시나리오를 중심으로 상황별 응답 거부 접근법을 제안한다.
    첫째, Universal Domain Adaptation (UniDA)은 미지의 외부 도메인에서 유입되는 입력에 대해 적응(adaptation)할지 혹은 응답을 거부할지를 판단해야 하는 문제다.
    본 연구는 자연어처리 분야 최초의 UniDA 통합 평가 환경을 구축하고, 기존 연구들이 해당 환경에서 어떻게 동작하는지 체계적으로 분석하였다.
    둘째로, 사용자로부터 모호한(ambiguous) 질의가 제공되었을 때를 탐구한다.
    모호한 질의는 하나의 질문이 복수의 유효한 해석을 가질 수 있기 때문에, 사용자 의도에 벗어날 수 있는 임의의 답변을 제공하는 것 보다 사용자로부터 명확한 의도를 확인하는 것이 중요하다.
    본 연구는 질의 내 모호성을 탐지하고 명확화(clarification) 요청을 생성함으로써 사용자의 의도에 벗어나는 딥변을 방지하는 학습 프레임워크를 제안한다.
    마지막으로, 외부 지식을 활용한 질의응답 시나리오를 다룬다.
    이런 상황에서 모델은 입력으로 제공된 외부 지식과 사전학습된 모델의 지식을 모두 사용할 수 있다.
    사용자로부터 주어진 질의에 답변하기 위한 명확한 근거가 부재한 경우 응답을 거부할 수 있도록, 외부 지식과 사전학습된 지식에 동적으로 가중치는 부여하는 디코딩 방법론을 제안한다.

    다양한 실험을 통해, 본 논문은 상황에 따른 답변 거부 메커니즘이 언어모델의 신뢰성, 강건성 및 안전성을 크게 향상시킴을 보였다.
    종합하면, 본 연구는 언어모델의 응답 거부 능력 향상을 위한 포괄적 프레임워크를 제시하며, 실세계 환경에서 신뢰 가능하고 책임 있는 언어모델의 활용 기반을 마련한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Language models (LMs) have achieved remarkable advances in natural language processing (NLP), yet their reliability remains fundamentally limited by the tendency to frequently generate inaccurate, ambiguous, or misaligned outputs.
    This dissertation addresses this challenge through the lens of abstention—the deliberate decision of a model to withhold unreliable or unintended responses.
    Specifically, we investigate task-specific abstention mechanisms, as the reliability requirements vary across different contexts and scenarios
    Since unintended or unreliable outputs should not be directly delivered to users, abstention mechanisms should be carefully designed to address distinct reliability challenged associated with each task.


    To this end, this dissertation explores three representative tasks, each illustrating distinct reliability challenges and motivating task-specific abstention.
    The first explores Universal Domain Adaptation (UniDA), where the challenge lies in determining whether test inputs from unseen domain should be adapted or abstained from.
    We construct the first unified benchmark of UniDA for NLP and systematically evaluate how existing approaches handle unpredictable out-of-distribution inputs.
    The second focuses on handling ambiguous user queries, where a single query may pose multiple valid interpretations.
    We introduce Alignment with Perceived Ambiguity (APA), a training framework which enables the models to detect ambiguity within the given query and abstain by generating clarification requests, thereby avoiding misinterpretation of user intent.
    Finally, we focus on contextual question-answering, where additional external knowledge source can supplement the model's pre-trained parametric knowledge.
    We propose Contrastive Decoding with Abstention (CDA), a training-free decoding method that dynamically attends to both knowledge sources and abstains when reliable grounding is unavailable.


    Through extensive experiments, we demonstrate that principled, task-specific abstention mechanisms significantly enhance the reliability, robustness, and safety of language models.
    Taken together, this dissertation contributes a comprehensive framework for abstention in LMs, laying the groundwork for their responsible and trustworthy deployment in real-world applications.
    번역하기

    Language models (LMs) have achieved remarkable advances in natural language processing (NLP), yet their reliability remains fundamentally limited by the tendency to frequently generate inaccurate, ambiguous, or misaligned outputs. This dissertation a...

    Language models (LMs) have achieved remarkable advances in natural language processing (NLP), yet their reliability remains fundamentally limited by the tendency to frequently generate inaccurate, ambiguous, or misaligned outputs.
    This dissertation addresses this challenge through the lens of abstention—the deliberate decision of a model to withhold unreliable or unintended responses.
    Specifically, we investigate task-specific abstention mechanisms, as the reliability requirements vary across different contexts and scenarios
    Since unintended or unreliable outputs should not be directly delivered to users, abstention mechanisms should be carefully designed to address distinct reliability challenged associated with each task.


    To this end, this dissertation explores three representative tasks, each illustrating distinct reliability challenges and motivating task-specific abstention.
    The first explores Universal Domain Adaptation (UniDA), where the challenge lies in determining whether test inputs from unseen domain should be adapted or abstained from.
    We construct the first unified benchmark of UniDA for NLP and systematically evaluate how existing approaches handle unpredictable out-of-distribution inputs.
    The second focuses on handling ambiguous user queries, where a single query may pose multiple valid interpretations.
    We introduce Alignment with Perceived Ambiguity (APA), a training framework which enables the models to detect ambiguity within the given query and abstain by generating clarification requests, thereby avoiding misinterpretation of user intent.
    Finally, we focus on contextual question-answering, where additional external knowledge source can supplement the model's pre-trained parametric knowledge.
    We propose Contrastive Decoding with Abstention (CDA), a training-free decoding method that dynamically attends to both knowledge sources and abstains when reliable grounding is unavailable.


    Through extensive experiments, we demonstrate that principled, task-specific abstention mechanisms significantly enhance the reliability, robustness, and safety of language models.
    Taken together, this dissertation contributes a comprehensive framework for abstention in LMs, laying the groundwork for their responsible and trustworthy deployment in real-world applications.

    더보기

    목차 (Table of Contents)

    • Chapter 1 Introduction 1
    • 1.1 Abstention in Language Models 1
    • 1.2 Dissertation Outline 6
    • Chapter 2 Background 8
    • 2.1 Problem Formulation 8
    • Chapter 1 Introduction 1
    • 1.1 Abstention in Language Models 1
    • 1.2 Dissertation Outline 6
    • Chapter 2 Background 8
    • 2.1 Problem Formulation 8
    • 2.1.1 Query Perspective 10
    • 2.1.2 Model Capability Perspective 11
    • 2.1.3 Human Value Alignment Perspective 12
    • 2.1.4 When to Abstain 13
    • 2.1.5 How to Abstain 14
    • 2.2 Taxonomy of Abstention in Language Models 15
    • 2.2.1 Inference-based Methods 16
    • 2.2.2 Training-based Methods 19
    • 2.3 Scope of the Dissertation 20
    • Chapter 3 Related Work 23
    • 3.1 Domain Adaptation (DA) 23
    • 3.2 Out-of-Distribution (OOD) Detection 24
    • 3.3 Ambiguity in NLP 25
    • 3.4 Human Preference Alignment 25
    • 3.4.1 Alignment Training of LLMs 25
    • 3.4.2 Quality-based Data Selection for Alignment 26
    • 3.5 Contrastive Decoding 27
    • 3.6 Abstention in LLMs 27
    • Chapter 4 Abstention in Distributional Shifts 29
    • 4.1 Introduction 29
    • 4.2 Problem Formulation 32
    • 4.2.1 Universal Domain Adaptation (UniDA) 32
    • 4.2.2 Challenges in UniDA 33
    • 4.3 Dataset Design 34
    • 4.3.1 Quantifying Different Shifts 34
    • 4.3.2 Implementation of Different Shifts 35
    • 4.3.3 Dataset Details 36
    • 4.3.4 Dataset Analysis 37
    • 4.4 Experimental Setting 39
    • 4.4.1 Compared Methods 39
    • 4.4.2 Thresholding Method 40
    • 4.4.3 Evaluation Metric 41
    • 4.4.4 Implementation Details 41
    • 4.5 Experimental Results 42
    • 4.5.1 Overview 42
    • 4.5.2 Detailed Results 43
    • 4.5.3 Impact of Threshold Values 46
    • 4.5.4 Impact of Different Scoring Functions 46
    • 4.5.5 Receiver Operating Characteristic (ROC) Curve 47
    • 4.6 Conclusion and Future Work 48
    • Chapter 5 Abstention for Ambiguous Inputs 51
    • 5.1 Introduction 51
    • 5.2 Alignment with Perceived Ambiguity (Apa) 55
    • 5.2.1 Initial Prediction Assessment 57
    • 5.2.2 Perceived Ambiguity Detection 57
    • 5.2.3 Response Construction 58
    • 5.2.4 Supervised Fine-Tuning (SFT) 59
    • 5.3 Experimental Setting 60
    • 5.3.1 Datasets 60
    • 5.3.2 Baselines 61
    • 5.3.3 Evaluation Metrics 64
    • 5.3.4 Implementation Details 65
    • 5.4 Experimental Results 65
    • 5.5 Ablation Study 67
    • 5.5.1 Analysis on Sample-level Misalignment 68
    • 5.5.2 The Effect of Threshold Values 69
    • 5.5.3 Impact of Infogain for Data Selection 70
    • 5.5.4 Comparison with State-of-the-Art Foundation Models 71
    • 5.6 Case Study 72
    • 5.6.1 Qualitative Analysis of Disambiguation and Clarification 72
    • 5.6.2 Failure Cases Before Alignment 73
    • 5.6.3 Failure Cases of Clarification Request Generation 73
    • 5.7 Conclusion 74
    • Chapter 6 Abstention in the Absence of Relevant Knowledge 75
    • 6.1 Introduction 75
    • 6.2 Dataset Design for Controlled Analysis 79
    • 6.2.1 Problem Formulation 79
    • 6.2.2 Initial Dataset Construction 80
    • 6.2.3 Parametric Knowledge Estimation 81
    • 6.2.4 Contextual Knowledge Estimation 81
    • 6.2.5 Final Dataset Construction 83
    • 6.3 Contrastive Decoding with Abstention 83
    • 6.3.1 Preliminary 84
    • 6.3.2 Incorporating Abstention 84
    • 6.3.3 Knowledge Relevance Assessment via Uncertainty Calibration 85
    • 6.3.4 CDA with Momentum (CDA-m) 86
    • 6.4 Experiments 87
    • 6.4.1 Experimental Setting 87
    • 6.4.2 Evaluation Metric 88
    • 6.4.3 Baselines 89
    • 6.4.4 Main Results 93
    • 6.5 Ablation Study 93
    • 6.5.1 Analysis of Different Scenarios 94
    • 6.5.2 Ablation on Momentum Weight 95
    • 6.5.3 Effect of Calibration 98
    • 6.5.4 Comparison with Training-based Methods 98
    • 6.5.5 Evaluation on RAG Setting 100
    • 6.5.6 Comparison with State-of-the-Art Foundation Models 102
    • 6.5.7 Output Distribution Analysis 104
    • 6.5.8 Computation Cost Analysis 105
    • 6.6 Conclusion 106
    • Chapter 7 Conclusion 107
    • 7.1 Summary 107
    • 7.2 Limitations and Future Work 109
    • Bibliography 111
    • 초록 149
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼