RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    어텐션 분포 정렬을 통한 생성형 언어 모델의 사회적 편향 완화 기법 연구 = Aligning Attention Distributions to Mitigate Social Bias in Generative Language Models

    한글로보기

    https://www.riss.kr/link?id=T17313619

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Large language models (LLMs) frequently generate outputs that reflect social biases, leading to growing concerns about fairness and potential harm. This thesis presents KLAAD (KL-Attention Alignment Debiasing), an attention-based debiasing framework that encourages alignment between the attention distributions of stereotypical and anti-stereotypical sentence pairs, without directly modifying the model's parameters. The proposed method incorporates a composite loss function--comprising Cross-Entropy, KL divergence, and Triplet losses--to promote consistent attention across biased and unbiased contexts while preserving the model's fluency and coherence. Experimental evaluations conducted on the BBQ and BOLD benchmarks demonstrate that KLAAD effectively mitigates bias with minimal impact on language modeling performance. These results suggest that guiding attention alignment offers a principled and effective approach for bias reduction in generative language models.
    번역하기

    Large language models (LLMs) frequently generate outputs that reflect social biases, leading to growing concerns about fairness and potential harm. This thesis presents KLAAD (KL-Attention Alignment Debiasing), an attention-based debiasing framework t...

    Large language models (LLMs) frequently generate outputs that reflect social biases, leading to growing concerns about fairness and potential harm. This thesis presents KLAAD (KL-Attention Alignment Debiasing), an attention-based debiasing framework that encourages alignment between the attention distributions of stereotypical and anti-stereotypical sentence pairs, without directly modifying the model's parameters. The proposed method incorporates a composite loss function--comprising Cross-Entropy, KL divergence, and Triplet losses--to promote consistent attention across biased and unbiased contexts while preserving the model's fluency and coherence. Experimental evaluations conducted on the BBQ and BOLD benchmarks demonstrate that KLAAD effectively mitigates bias with minimal impact on language modeling performance. These results suggest that guiding attention alignment offers a principled and effective approach for bias reduction in generative language models.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    대규모 언어 모델(Large Language Models, LLMs)은 생성 문장에서 사회적 편향을 내포하는 경향이 있으며, 이는 공정성과 잠재적 피해 측면에서 윤리적인 문제를 야기한다. 본 논문에서는 이러한 편향 문제를 완화하기 위해 어텐션 기반의 편향 완화 프레임워크인 KLAAD (KL-Attention Alignment Debiasing)를 제안한다. 이 방법은 고정관념(stereotype) 문장과 반고정관념(anti-stereotype) 문장 간의 어텐션 분포를 정렬(alignment)하도록 유도함으로써, 모델의 가중치를 직접 변경하지 않고도 편향을 줄이는 것을 목표로 한다. 이를 위해 Cross-Entropy, KL divergence, Triplet loss로 구성된 복합 손실 함수를 설계하여, 사회적 편향을 내포한 문맥과 그렇지 않은 문맥 모두에서 일관된 어텐션 분포가 유지되도록 모델을 학습시킨다. BBQ 및 BOLD 벤치마크에서 수행한 실험 결과, 제안하는 KLAAD는 모델의 언어 생성 성능을 크게 저하시키지 않으면서도 편향을 효과적으로 줄이는 것으로 나타났다. 본 연구는 어텐션 정렬 기반의 접근법이 생성형 언어 모델에서 사회적 편향을 완화하는 실질적인 방법이 될 수 있음을 시사한다.
    번역하기

    대규모 언어 모델(Large Language Models, LLMs)은 생성 문장에서 사회적 편향을 내포하는 경향이 있으며, 이는 공정성과 잠재적 피해 측면에서 윤리적인 문제를 야기한다. 본 논문에서는 이러한 편...

    대규모 언어 모델(Large Language Models, LLMs)은 생성 문장에서 사회적 편향을 내포하는 경향이 있으며, 이는 공정성과 잠재적 피해 측면에서 윤리적인 문제를 야기한다. 본 논문에서는 이러한 편향 문제를 완화하기 위해 어텐션 기반의 편향 완화 프레임워크인 KLAAD (KL-Attention Alignment Debiasing)를 제안한다. 이 방법은 고정관념(stereotype) 문장과 반고정관념(anti-stereotype) 문장 간의 어텐션 분포를 정렬(alignment)하도록 유도함으로써, 모델의 가중치를 직접 변경하지 않고도 편향을 줄이는 것을 목표로 한다. 이를 위해 Cross-Entropy, KL divergence, Triplet loss로 구성된 복합 손실 함수를 설계하여, 사회적 편향을 내포한 문맥과 그렇지 않은 문맥 모두에서 일관된 어텐션 분포가 유지되도록 모델을 학습시킨다. BBQ 및 BOLD 벤치마크에서 수행한 실험 결과, 제안하는 KLAAD는 모델의 언어 생성 성능을 크게 저하시키지 않으면서도 편향을 효과적으로 줄이는 것으로 나타났다. 본 연구는 어텐션 정렬 기반의 접근법이 생성형 언어 모델에서 사회적 편향을 완화하는 실질적인 방법이 될 수 있음을 시사한다.

    더보기

    목차 (Table of Contents)

    • 제 1 장 서론 1
    • 제 1 절 연구의 목적과 배경 1
    • 제 2 절 연구의 내용 3
    • 제 2 장 문헌 고찰 5
    • 제 1 장 서론 1
    • 제 1 절 연구의 목적과 배경 1
    • 제 2 절 연구의 내용 3
    • 제 2 장 문헌 고찰 5
    • 제 1 절 편향 완화 방법론 5
    • 제 2 절 어텐션 메커니즘 6
    • 제 3 장 연구 방법 8
    • 제 1 절 개요 8
    • 제 2 절 데이터셋 8
    • 제 3 절 목적 함수 11
    • 제 4 장 연구 결과 및 분석 14
    • 제 1 절 실험 설정 14
    • 1. 모델 구성 및 학습 설정 14
    • 2. 비교 대상 편향 완화 방법론 14
    • 3. 평가 데이터셋 및 지표 15
    • 제 2 절 실험 결과 및 분석 19
    • 1. 어텐션 분포 변화 분석 19
    • 2. 성능 평가 결과 1 - BBQ 21
    • 3. 성능 평가 결과 2 - BOLD 22
    • 4. 성능 평가 결과 3 - CrowS-Pairs 25
    • 5. 하이퍼파라미터 조정 실험 33
    • 6. 소거 실험 34
    • 제 5 장 결론 37
    • 참고문헌 39
    • Abstract 46
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼