본 학위논문은 머신러닝 모델의 신뢰성을 향상시키기 위한 방법론으로, 의미 정합성 개선과 공정한 잠재 공간의 구성을 중심으로 연구를 수행하였다. 신뢰할 수 있는 머신러닝 시스템은 인...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
본 학위논문은 머신러닝 모델의 신뢰성을 향상시키기 위한 방법론으로, 의미 정합성 개선과 공정한 잠재 공간의 구성을 중심으로 연구를 수행하였다. 신뢰할 수 있는 머신러닝 시스템은 인...
본 학위논문은 머신러닝 모델의 신뢰성을 향상시키기 위한 방법론으로, 의미 정합성 개선과 공정한 잠재 공간의 구성을 중심으로 연구를 수행하였다. 신뢰할 수 있는 머신러닝 시스템은 인간의 직관에 부합하는 예측과 설명을 제공해야 하며, 다양한 조건에서도 안정적으로 작동하고, 민감한 속성에 따른 편향을 완화함으로써 공정성을 확보해야 한다. 이를 달성하기 위해 본 논문은 세 가지 상호 연관된 연구 방향에 초점을 맞춘다.
첫째, 신뢰 가능하고 해석 가능한 설명 공간을 구축하기 위한 분포 기반 프로토타입 기법을 제안한다. 기존의 프로토타입 기반 설명 방식은 유사한 입력에 대해 서로 다른 거리를 제공하는 등 설명의 일관성이 부족한 경우가 많다. 본 연구에서는 분포 임베딩과 잠재 공간 정규화를 통해 유사한 입력이 유사한 잠재 표현으로 매핑되도록 하여, 설명의 안정성과 해석 가능성을 향상시킨다. 또한 입력과 프로토타입 간의 특징 변화 과정을 시각적으로 제시함으로써 사용자가 설명을 보다 직관적으로 이해할 수 있도록 하였다.
둘째, 텍스트-이미지 생성 확산 모델에서 발생하는 의미 왜곡과 데이터 기억 현상을 완화하기 위한 의미 정합성 향상 기법을 제안한다. 확산 기반 생성 모델은 학습 데이터를 그대로 재현하는 경향이 있어 개인정보 유출 및 저작권 침해와 같은 문제가 발생할 수 있다. 이를 해결하기 위해, 의미 있는 토큰에 주의를 집중하도록 cross-attention map을 조정하는 새로운 추론 기반 완화 기법을 제안하였다. 이 방법은 의미 정합성을 유지하면서도 불필요한 기억을 줄이는 데 효과적임을 실험을 통해 확인하였다.
셋째, 민감한 속성과 예측에 필요한 정보를 분리하여 공정한 잠재 표현 공간을 구성하는 방법론을 개발하였다. 가역 신경망과 사전 학습된 생성 모델을 활용하여 민감 속성과 주요 예측 요인이 서로 다른 차원에 표현되도록 설계하였으며, 이를 통해 편향 없는 반사실 예제를 생성하고, 모델의 공정성을 보다 명확하게 설명할 수 있도록 하였다.
종합적으로, 본 논문은 의미 정합성, 해석 가능성, 공정성을 개선함으로써 머신러닝 모델의 신뢰성을 높이기 위한 다양한 방법론을 제시하였다. 제안한 기법들은 실제 응용에 적용 가능한 형태로 설계되었으며, 설명의 일관성 확보, 생성 모델의 기억 완화, 공정한 표현 학습 등 신뢰 가능한 인공지능 시스템 구축을 위한 기반을 제공한다.
다국어 초록 (Multilingual Abstract)
This dissertation addresses the critical challenge of enhancing the reliability of machine learning models by improving semantic alignment and constructing fair latent spaces. Reliable machine learning systems must produce predictions and explanations...
This dissertation addresses the critical challenge of enhancing the reliability of machine learning models by improving semantic alignment and constructing fair latent spaces. Reliable machine learning systems must produce predictions and explanations that align closely with human intuition, remain stable under varying conditions, and ensure fairness by mitigating biases associated with sensitive attributes. To meet these requirements, this work focuses on three interconnected directions.
First, I propose distributional prototypical methods for constructing reliable and interpretable explanation spaces. Traditional prototype-based explanations often suffer from inconsistencies, where similar inputs can yield dissimilar explanatory distances. By incorporating distributional embeddings and latent space regularization, the proposed methods ensure that similar inputs are mapped to similar latent representations, thereby improving stability and interpretability. In addition, I introduce transitional visualization techniques that help users intuitively understand how features evolve between inputs and prototypes. These methods help reduce the semantic gap between human perception and model explanations, while also providing robustness against input perturbations such as noise and compression.
Second, I explore techniques to improve semantic alignment in generative models, particularly in text-to-image diffusion models. These models are susceptible to memorizing training data, raising concerns about privacy and copyright infringement. To address this, I propose a novel inference-time mitigation technique that adjusts cross-attention maps to focus on semantically meaningful tokens. This approach effectively reduces verbatim memorization while preserving prompt fidelity, thereby enhancing semantic alignment and mitigating risks associated with model misuse.
Third, I develop methods for constructing fair latent spaces by disentangling sensitive attributes from task-relevant features. Leveraging invertible neural networks and pre-trained generative models, I construct structured latent representations in which sensitive and decision-relevant factors occupy distinct dimensions. This disentanglement enables the generation of unbiased counterfactuals and provides interpretable assessments of model fairness across diverse demographic groups.
In conclusion, this dissertation presents a set of methodologies aimed at improving the reliability of machine learning systems through enhanced semantic alignment, interpretability, and fairness. These contributions advance the understanding and development of trustworthy AI models by offering practical solutions for reducing semantic inconsistencies, mitigating memorization in generative models, and promoting fairness through disentangled representations. The proposed approaches lay a foundation for building more transparent and responsible machine learning systems applicable to a wide range of real-world domains.
목차 (Table of Contents)