언어모델은 자연어 처리 분야에서 눈부신 발전을 이루었으나, 부정확하거나 사용자 의도에 부합하지 않는 응답을 빈번히 생성하는 경향으로 인해 여전히 신뢰성 측면에서 근본적인 한계를 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17452093
서울 : 서울대학교 대학원, 2026
학위논문(박사) -- 서울대학교 대학원 , 협동과정인공지능전공 , 2026. 2
2026
영어
006.3
서울
xv, 150 ; 26 cm
지도교수: 이상구
I804:11032-000000194120
0
상세조회0
다운로드언어모델은 자연어 처리 분야에서 눈부신 발전을 이루었으나, 부정확하거나 사용자 의도에 부합하지 않는 응답을 빈번히 생성하는 경향으로 인해 여전히 신뢰성 측면에서 근본적인 한계를 ...
언어모델은 자연어 처리 분야에서 눈부신 발전을 이루었으나, 부정확하거나 사용자 의도에 부합하지 않는 응답을 빈번히 생성하는 경향으로 인해 여전히 신뢰성 측면에서 근본적인 한계를 지닌다.
본 논문은 이러한 문제를 해결하기 위해, 모델이 신뢰하기 어렵거나 의도에서 벗어난 응답을 명시적으로 거부(abstain)하는 관점에서 접근한다.
특히, 응답의 신뢰성에 대한 조건과 요구사항은 주어진 상황에 따라 크게 상이함으로, 본 연구는 주어진 시나리오의 특성에 적합하게 설계된 상황별 응답 거부(task-specific abstention) 접근법을 탐구한다.
이에 따라, 본 논문은 세 가지 대표적인 시나리오를 중심으로 상황별 응답 거부 접근법을 제안한다.
첫째, Universal Domain Adaptation (UniDA)은 미지의 외부 도메인에서 유입되는 입력에 대해 적응(adaptation)할지 혹은 응답을 거부할지를 판단해야 하는 문제다.
본 연구는 자연어처리 분야 최초의 UniDA 통합 평가 환경을 구축하고, 기존 연구들이 해당 환경에서 어떻게 동작하는지 체계적으로 분석하였다.
둘째로, 사용자로부터 모호한(ambiguous) 질의가 제공되었을 때를 탐구한다.
모호한 질의는 하나의 질문이 복수의 유효한 해석을 가질 수 있기 때문에, 사용자 의도에 벗어날 수 있는 임의의 답변을 제공하는 것 보다 사용자로부터 명확한 의도를 확인하는 것이 중요하다.
본 연구는 질의 내 모호성을 탐지하고 명확화(clarification) 요청을 생성함으로써 사용자의 의도에 벗어나는 딥변을 방지하는 학습 프레임워크를 제안한다.
마지막으로, 외부 지식을 활용한 질의응답 시나리오를 다룬다.
이런 상황에서 모델은 입력으로 제공된 외부 지식과 사전학습된 모델의 지식을 모두 사용할 수 있다.
사용자로부터 주어진 질의에 답변하기 위한 명확한 근거가 부재한 경우 응답을 거부할 수 있도록, 외부 지식과 사전학습된 지식에 동적으로 가중치는 부여하는 디코딩 방법론을 제안한다.
다양한 실험을 통해, 본 논문은 상황에 따른 답변 거부 메커니즘이 언어모델의 신뢰성, 강건성 및 안전성을 크게 향상시킴을 보였다.
종합하면, 본 연구는 언어모델의 응답 거부 능력 향상을 위한 포괄적 프레임워크를 제시하며, 실세계 환경에서 신뢰 가능하고 책임 있는 언어모델의 활용 기반을 마련한다.
다국어 초록 (Multilingual Abstract)
Language models (LMs) have achieved remarkable advances in natural language processing (NLP), yet their reliability remains fundamentally limited by the tendency to frequently generate inaccurate, ambiguous, or misaligned outputs. This dissertation a...
Language models (LMs) have achieved remarkable advances in natural language processing (NLP), yet their reliability remains fundamentally limited by the tendency to frequently generate inaccurate, ambiguous, or misaligned outputs.
This dissertation addresses this challenge through the lens of abstention—the deliberate decision of a model to withhold unreliable or unintended responses.
Specifically, we investigate task-specific abstention mechanisms, as the reliability requirements vary across different contexts and scenarios
Since unintended or unreliable outputs should not be directly delivered to users, abstention mechanisms should be carefully designed to address distinct reliability challenged associated with each task.
To this end, this dissertation explores three representative tasks, each illustrating distinct reliability challenges and motivating task-specific abstention.
The first explores Universal Domain Adaptation (UniDA), where the challenge lies in determining whether test inputs from unseen domain should be adapted or abstained from.
We construct the first unified benchmark of UniDA for NLP and systematically evaluate how existing approaches handle unpredictable out-of-distribution inputs.
The second focuses on handling ambiguous user queries, where a single query may pose multiple valid interpretations.
We introduce Alignment with Perceived Ambiguity (APA), a training framework which enables the models to detect ambiguity within the given query and abstain by generating clarification requests, thereby avoiding misinterpretation of user intent.
Finally, we focus on contextual question-answering, where additional external knowledge source can supplement the model's pre-trained parametric knowledge.
We propose Contrastive Decoding with Abstention (CDA), a training-free decoding method that dynamically attends to both knowledge sources and abstains when reliable grounding is unavailable.
Through extensive experiments, we demonstrate that principled, task-specific abstention mechanisms significantly enhance the reliability, robustness, and safety of language models.
Taken together, this dissertation contributes a comprehensive framework for abstention in LMs, laying the groundwork for their responsible and trustworthy deployment in real-world applications.
목차 (Table of Contents)