RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Evaluating and Interpreting Implicit Social Bias in Large Language Models via Defeasible Reasoning = 폐기 가능 추론을 통한 거대 언어 모델의 암묵적 사회 편향 평가와 해석 연구

    한글로보기

    https://www.riss.kr/link?id=T17450270

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    While large language models (LLMs) have demonstrated remarkable capabilities, they often exhibit significant social biases. However, conventional evaluation frameworks are largely limited to detecting a narrow set of explicit social stereotypes, failing to capture the subtle and implicit biases inherent in modern models. To address this, this study redefines implicit bias as behavioral inconsistencies in model decision-making and presents a comprehensive framework for diagnosing and interpreting these implicit social biases through the lens of defeasible reasoning, a non-monotonic reasoning where inferences are subject to revision based on new information. By adopting a defeasible reasoning framework, this research examines how demographic cues disproportionately influence model reasoning.
    For this purpose, this study proposes an evaluation framework comprising a dedicated dataset, curated from benign social commonsense and natural language inference corpora and adapted for five bias axes, alongside new metrics designed to quantify prediction inconsistencies. Empirical evaluation of ten instruction-tuned LLMs reveals that models exhibit systematic biases that align with real-world social hierarchies, where introducing cues to particular demographics triggers disproportionate shifts in their decisions. Notably, explicitly engaging in step-by-step reasoning further amplifies these biases, triggering the models’ latent stereotypes in intermediate reasoning steps.
    In addition to empirical analysis, this study applies mechanistic interpretability techniques, specifically attribution patching and edge ablation, to localize the structural origins of these biases as a case study. Moving beyond prior interpretability research limited to explicit stereotypes, this analysis identifies that demographic-sensitive behavior difference is mediated by a remarkably sparse set of critical paths within a model. By shifting the focus from what models explicitly say to how they implicitly reason, this study provides a starting point for more systematic diagnosis and analysis of reasoning behavior and fairness in LLMs.
    번역하기

    While large language models (LLMs) have demonstrated remarkable capabilities, they often exhibit significant social biases. However, conventional evaluation frameworks are largely limited to detecting a narrow set of explicit social stereotypes, faili...

    While large language models (LLMs) have demonstrated remarkable capabilities, they often exhibit significant social biases. However, conventional evaluation frameworks are largely limited to detecting a narrow set of explicit social stereotypes, failing to capture the subtle and implicit biases inherent in modern models. To address this, this study redefines implicit bias as behavioral inconsistencies in model decision-making and presents a comprehensive framework for diagnosing and interpreting these implicit social biases through the lens of defeasible reasoning, a non-monotonic reasoning where inferences are subject to revision based on new information. By adopting a defeasible reasoning framework, this research examines how demographic cues disproportionately influence model reasoning.
    For this purpose, this study proposes an evaluation framework comprising a dedicated dataset, curated from benign social commonsense and natural language inference corpora and adapted for five bias axes, alongside new metrics designed to quantify prediction inconsistencies. Empirical evaluation of ten instruction-tuned LLMs reveals that models exhibit systematic biases that align with real-world social hierarchies, where introducing cues to particular demographics triggers disproportionate shifts in their decisions. Notably, explicitly engaging in step-by-step reasoning further amplifies these biases, triggering the models’ latent stereotypes in intermediate reasoning steps.
    In addition to empirical analysis, this study applies mechanistic interpretability techniques, specifically attribution patching and edge ablation, to localize the structural origins of these biases as a case study. Moving beyond prior interpretability research limited to explicit stereotypes, this analysis identifies that demographic-sensitive behavior difference is mediated by a remarkably sparse set of critical paths within a model. By shifting the focus from what models explicitly say to how they implicitly reason, this study provides a starting point for more systematic diagnosis and analysis of reasoning behavior and fairness in LLMs.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    거대 언어 모델은 뛰어난 언어 능력을 보이는 것으로 평가되나 동시에 상당한 수준의 편향성을 드러내기도 한다. 그러나 기존의 평가 프레임워크는 주로 일부의 명시적 고정관념의 탐지에 국한되어 있어, 현대 모델 내에 존재하는 보다 암묵적인 편향을 포착하는 데 한계가 있다. 본 연구는 이러한 문제를 해결하기 위하여 암묵적 편향을 모델 의사결정의 행동적 비일관성으로 재정의하고, 새로운 정보에 따라 추론 결과가 수정되는 비단조적 추론인 폐기 가능 추론을 활용하여, 인구통계학적 정보 단서가 모델의 추론 및 논리적 의사결정에 얼마나 불균형적인 영향을 미치는지의 관점에서 이러한 편향을 진단하고 해석할 수 있는 프레임워크를 제안한다.
    이를 위해 본 연구는 사회적 상식 추론 및 자연어 추론 코퍼스를 재구성하여 제작한 전용 평가 데이터셋을 구축, 예측의 비일관성을 정량화하는 새로운 평가 지표를 포함한 프레임워크를 도입한다. 이를 바탕으로 10종의 거대 언어 모델을 대상으로 평가를 수행하였으며 그 결과 모델들이 실제 사회적 차별 구조와 일치하는 양상으로 나타나는 보이는 체계적인 의사결정 편향을 보임을 확인하였다. 특정 집단에 대한 단서가 주어질 때 모델의 추론 결과가 덜 혹은 더 이질적으로 변화하는 양상이 관찰되었으며, 특히 명시적인 단계별 추론을 적용할 때 모델이 잠재된 고정관념을 바탕으로 논리를 전개함으로써 오히려 편향이 증폭되는 현상이 나타났다.
    모델에 대한 평가를 넘어 본 연구는 어트리뷰션 패칭과 엣지 제거 등 기계론적 해석 가능성 기법을 적용하여 이러한 편향의 메커니즘을 확인하고자 하였다. 명시적인 소수의 고정관념에만 적용되었던 기존의 해석적 연구들과 달리 본 분석은 인구통계학적 정보에 민감하게 반응하는 모델의 행동적 차이를 모델 내부적 관점에서 분석하였으며, 모델 내부의 희소한 핵심 경로들에 의해 그러한 편향이 매개된다는 점을 식별하였다.
    본 연구는 모델이 명시적으로 생성하는 텍스트에 대한 감지에서 나아가 모델의 의사결정 과정과 추론으로 분석의 초점을 전환함으로써 거대 언어 모델의 추론 행동과 편향 및 공정성에 대한 보다 체계적인 진단 및 분석의 토대를 제공한다.
    번역하기

    거대 언어 모델은 뛰어난 언어 능력을 보이는 것으로 평가되나 동시에 상당한 수준의 편향성을 드러내기도 한다. 그러나 기존의 평가 프레임워크는 주로 일부의 명시적 고정관념의 탐지에 ...

    거대 언어 모델은 뛰어난 언어 능력을 보이는 것으로 평가되나 동시에 상당한 수준의 편향성을 드러내기도 한다. 그러나 기존의 평가 프레임워크는 주로 일부의 명시적 고정관념의 탐지에 국한되어 있어, 현대 모델 내에 존재하는 보다 암묵적인 편향을 포착하는 데 한계가 있다. 본 연구는 이러한 문제를 해결하기 위하여 암묵적 편향을 모델 의사결정의 행동적 비일관성으로 재정의하고, 새로운 정보에 따라 추론 결과가 수정되는 비단조적 추론인 폐기 가능 추론을 활용하여, 인구통계학적 정보 단서가 모델의 추론 및 논리적 의사결정에 얼마나 불균형적인 영향을 미치는지의 관점에서 이러한 편향을 진단하고 해석할 수 있는 프레임워크를 제안한다.
    이를 위해 본 연구는 사회적 상식 추론 및 자연어 추론 코퍼스를 재구성하여 제작한 전용 평가 데이터셋을 구축, 예측의 비일관성을 정량화하는 새로운 평가 지표를 포함한 프레임워크를 도입한다. 이를 바탕으로 10종의 거대 언어 모델을 대상으로 평가를 수행하였으며 그 결과 모델들이 실제 사회적 차별 구조와 일치하는 양상으로 나타나는 보이는 체계적인 의사결정 편향을 보임을 확인하였다. 특정 집단에 대한 단서가 주어질 때 모델의 추론 결과가 덜 혹은 더 이질적으로 변화하는 양상이 관찰되었으며, 특히 명시적인 단계별 추론을 적용할 때 모델이 잠재된 고정관념을 바탕으로 논리를 전개함으로써 오히려 편향이 증폭되는 현상이 나타났다.
    모델에 대한 평가를 넘어 본 연구는 어트리뷰션 패칭과 엣지 제거 등 기계론적 해석 가능성 기법을 적용하여 이러한 편향의 메커니즘을 확인하고자 하였다. 명시적인 소수의 고정관념에만 적용되었던 기존의 해석적 연구들과 달리 본 분석은 인구통계학적 정보에 민감하게 반응하는 모델의 행동적 차이를 모델 내부적 관점에서 분석하였으며, 모델 내부의 희소한 핵심 경로들에 의해 그러한 편향이 매개된다는 점을 식별하였다.
    본 연구는 모델이 명시적으로 생성하는 텍스트에 대한 감지에서 나아가 모델의 의사결정 과정과 추론으로 분석의 초점을 전환함으로써 거대 언어 모델의 추론 행동과 편향 및 공정성에 대한 보다 체계적인 진단 및 분석의 토대를 제공한다.

    더보기

    목차 (Table of Contents)

    • 1. Introduction 1
    • 2. Related Works 4
    • 2.1. Social Bias Evaluation 4
    • 2.1.1. Explicit Social Bias 4
    • 2.1.2. Implicit Social Bias 5
    • 1. Introduction 1
    • 2. Related Works 4
    • 2.1. Social Bias Evaluation 4
    • 2.1.1. Explicit Social Bias 4
    • 2.1.2. Implicit Social Bias 5
    • 2.2. Mechanistic Interpretability and Bias Interpretation 8
    • 2.2.1. Background of Mechanistic Interpretability 8
    • 2.2.2. Interpreting Social Bias 9
    • 3. Methodologies 12
    • 3.1. Task Formulation 12
    • 3.2. Evaluation Metrics 15
    • 3.2.1. Bias Score 16
    • 3.2.2. Base Mismatch Rate (BMR) 18
    • 3.3. Interpretation Methods 20
    • 3.3.1. Attribution Patching 20
    • 3.3.2. Edge Ablation 22
    • 4. Datasets 24
    • 4.1. Source Dataset 24
    • 4.1.1. SocialIQa 24
    • 4.1.2. ROC Stories 26
    • 4.1.3. δ-SNLI, δ-ATOMIC, δ-SOCIALCHEM 27
    • 4.2. Data Processing and Augmentation 30
    • 4.2.1. Normalization of Premise-Hypothesis Pairs 31
    • 4.2.2. Removal of Demographic Cues 31
    • 4.2.3. Construction of Demographic Update Sentences 32
    • 4.2.4. Base Update as a Reference Profile 34
    • 5. Experiments 36
    • 5.1. Bias Evaluation Experiment 36
    • 5.1.1. Models 36
    • 5.1.2. Prompt Style Variations 37
    • 5.1.3. Reasoning Mode Variations 38
    • 5.1.4. Generation Configuration 39
    • 5.2. Interpretation Experiments 40
    • 5.2.1. Model 41
    • 5.2.2. Datasets 42
    • 5.2.3. Prompts and Configurations 43
    • 6. Results and Analysis 44
    • 6.1. Evaluation Results and Analysis 44
    • 6.1.1. Bias Score Overview 44
    • 6.1.2. Analysis on Group-Level Response Patterns 48
    • 6.1.3. Analysis on Rationales 55
    • 6.2. Interpretation Results and Analysis 61
    • 6.2.1. Attribution Analysis 61
    • 6.2.2. Intervention Results and Analysis 64
    • 7. Conclusion 68
    • Bibliography 71
    • 국문초록 76
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼