RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Evaluating the Stability of Risky Decision-Making in Large Language Models : Cross-Task Consistency and Linguistic Framing Sensitivity Analysis = 거대 언어모델의 위험 의사결정 안정성 평가: 과제 간 일관성과 언어적 프레이밍 민감도 분석

    한글로보기

    https://www.riss.kr/link?id=T17450276

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Large language models (LLMs) are increasingly deployed in settings where their outputs influence human decision-making, including domains that involve risk, uncertainty, and probabilistic choices. As these models begin to function not only as information providers but also as decision-support agents, understanding the stability and reliability of their behavioral tendencies becomes essential.
    This thesis investigates two core questions: whether LLMs exhibit consistent patterns of risky decision-making across different tasks, and how strongly their choices shift under different linguistic framing conditions, with particular attention to differences between pretrained and instruction-tuned models. Using 11 models across pretrained and instruction-tuned variants, we evaluate risk-taking behavior on three psychological tasks—BART, CCT, and GDT—under Gain, Loss, and Neutral frames.
    Results reveal limited cross-task consistency overall: models that behave risk-seeking in one task often behave conservatively or unpredictably in another. Notably, pretrained models tend to exhibit more coherent risk-taking patterns across tasks, whereas instruction-tuned models show greater variability in cross-task behavior. Linguistic framing sensitivity also varies across training stages. Instruction-tuned models exhibit clearer gain–loss asymmetries and stronger responsiveness to linguistic framing, while pretrained models display weaker framing effects.
    Together, these findings indicate that instruction tuning improves responsiveness to linguistic cues but may come at the cost of cross-task behavioral stability. More broadly, they suggest that LLM-based decision-making is jointly shaped by task structure and linguistic formulation, underscoring the need for rigorous behavioral evaluation before integrating such models into high-stakes or decision-support contexts.
    번역하기

    Large language models (LLMs) are increasingly deployed in settings where their outputs influence human decision-making, including domains that involve risk, uncertainty, and probabilistic choices. As these models begin to function not only as informat...

    Large language models (LLMs) are increasingly deployed in settings where their outputs influence human decision-making, including domains that involve risk, uncertainty, and probabilistic choices. As these models begin to function not only as information providers but also as decision-support agents, understanding the stability and reliability of their behavioral tendencies becomes essential.
    This thesis investigates two core questions: whether LLMs exhibit consistent patterns of risky decision-making across different tasks, and how strongly their choices shift under different linguistic framing conditions, with particular attention to differences between pretrained and instruction-tuned models. Using 11 models across pretrained and instruction-tuned variants, we evaluate risk-taking behavior on three psychological tasks—BART, CCT, and GDT—under Gain, Loss, and Neutral frames.
    Results reveal limited cross-task consistency overall: models that behave risk-seeking in one task often behave conservatively or unpredictably in another. Notably, pretrained models tend to exhibit more coherent risk-taking patterns across tasks, whereas instruction-tuned models show greater variability in cross-task behavior. Linguistic framing sensitivity also varies across training stages. Instruction-tuned models exhibit clearer gain–loss asymmetries and stronger responsiveness to linguistic framing, while pretrained models display weaker framing effects.
    Together, these findings indicate that instruction tuning improves responsiveness to linguistic cues but may come at the cost of cross-task behavioral stability. More broadly, they suggest that LLM-based decision-making is jointly shaped by task structure and linguistic formulation, underscoring the need for rigorous behavioral evaluation before integrating such models into high-stakes or decision-support contexts.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    거대 언어모델(Large Language Models, LLMs)은 점차 단순한 텍스트 생성 도구를 넘어, 인간의 판단과 선택에 영향을 미치는 의사결정 지원 시스템으로 활용되고 있다. 특히 위험, 불확실성, 확률적 선택이 수반되는 영역에서 이러한 모델이 제시하는 판단은 실제 사용자에게 중대한 영향을 미칠 수 있다. 이에 따라, LLM이 위험 상황에서 얼마나 안정적이고 신뢰 가능한 행동 양상을 보이는지를 평가하는 것은 중요한 연구 과제가 되고 있다.
    본 논문은 LLM의 위험 의사결정 행동을 두 가지 관점에서 분석한다. 첫째, 모델이 서로 다른 의사결정 과제 간에 일관된 위험 선호 경향을 보이는지 여부이며, 둘째, 동일한 의사결정 구조 하에서 언어적 프레이밍 변화가 모델의 선택에 얼마나 큰 영향을 미치는지이다. 특히 사전학습(pretrained) 모델과 지시학습(instruction-tuned) 모델 간의 차이에 주목하여 분석을 수행하였다. 이를 위해 사전학습 및 지시학습 단계를 포함한 총 11개의 LLM을 대상으로, 풍선 아날로그 위험 과제(Balloon Analogue Risk Task, BART), 콜롬비아 카드 과제(Columbia Card Task, CCT), 주사위 게임 과제(Game of Dice Task, GDT)라는 세 가지 심리학적 위험 의사결정 과제를 적용하였다. 각 과제는 이득(gain), 손실(loss), 중립(neutral)의 세 가지 언어적 프레이밍 조건 하에서 반복 수행되었다.
    분석 결과, 전반적으로 과제 간 위험 의사결정 일관성은 낮은 것으로 나타났다. 한 과제에서 위험 선호적 행동을 보인 모델이 다른 과제에서는 보수적이거나 불안정한 선택을 보이는 경우가 빈번하게 관찰되었다. 특히 사전학습 모델은 과제 간 비교적 일관된 위험 선택 패턴을 보인 반면, 지시학습 모델은 과제에 따라 행동 변동성이 더 크게 나타났다. 한편 언어적 프레이밍 민감도 측면에서도 학습 단계에 따른 차이가 뚜렷하게 관찰되었다. 지시학습 모델은 이득–손실 프레이밍에 대해 상이한 반응을 보이며 언어적 변화에 민감하게 반응한 반면, 사전학습 모델은 프레이밍 효과가 상대적으로 약하게 나타났다.
    이러한 결과는 지시학습이 언어적 단서에 대한 반응성을 향상시키는 대신, 과제 전반에 걸친 행동 안정성을 저해할 수 있음을 시사한다. 더 나아가, LLM의 의사결정 행동은 과제 구조와 언어적 표현에 의해 공동으로 형성된다는 점을 보여주며, 고위험 환경이나 의사결정 지원 맥락에 LLM을 적용하기에 앞서 체계적인 행동 평가가 필요함을 강조한다.
    번역하기

    거대 언어모델(Large Language Models, LLMs)은 점차 단순한 텍스트 생성 도구를 넘어, 인간의 판단과 선택에 영향을 미치는 의사결정 지원 시스템으로 활용되고 있다. 특히 위험, 불확실성, 확률적 ...

    거대 언어모델(Large Language Models, LLMs)은 점차 단순한 텍스트 생성 도구를 넘어, 인간의 판단과 선택에 영향을 미치는 의사결정 지원 시스템으로 활용되고 있다. 특히 위험, 불확실성, 확률적 선택이 수반되는 영역에서 이러한 모델이 제시하는 판단은 실제 사용자에게 중대한 영향을 미칠 수 있다. 이에 따라, LLM이 위험 상황에서 얼마나 안정적이고 신뢰 가능한 행동 양상을 보이는지를 평가하는 것은 중요한 연구 과제가 되고 있다.
    본 논문은 LLM의 위험 의사결정 행동을 두 가지 관점에서 분석한다. 첫째, 모델이 서로 다른 의사결정 과제 간에 일관된 위험 선호 경향을 보이는지 여부이며, 둘째, 동일한 의사결정 구조 하에서 언어적 프레이밍 변화가 모델의 선택에 얼마나 큰 영향을 미치는지이다. 특히 사전학습(pretrained) 모델과 지시학습(instruction-tuned) 모델 간의 차이에 주목하여 분석을 수행하였다. 이를 위해 사전학습 및 지시학습 단계를 포함한 총 11개의 LLM을 대상으로, 풍선 아날로그 위험 과제(Balloon Analogue Risk Task, BART), 콜롬비아 카드 과제(Columbia Card Task, CCT), 주사위 게임 과제(Game of Dice Task, GDT)라는 세 가지 심리학적 위험 의사결정 과제를 적용하였다. 각 과제는 이득(gain), 손실(loss), 중립(neutral)의 세 가지 언어적 프레이밍 조건 하에서 반복 수행되었다.
    분석 결과, 전반적으로 과제 간 위험 의사결정 일관성은 낮은 것으로 나타났다. 한 과제에서 위험 선호적 행동을 보인 모델이 다른 과제에서는 보수적이거나 불안정한 선택을 보이는 경우가 빈번하게 관찰되었다. 특히 사전학습 모델은 과제 간 비교적 일관된 위험 선택 패턴을 보인 반면, 지시학습 모델은 과제에 따라 행동 변동성이 더 크게 나타났다. 한편 언어적 프레이밍 민감도 측면에서도 학습 단계에 따른 차이가 뚜렷하게 관찰되었다. 지시학습 모델은 이득–손실 프레이밍에 대해 상이한 반응을 보이며 언어적 변화에 민감하게 반응한 반면, 사전학습 모델은 프레이밍 효과가 상대적으로 약하게 나타났다.
    이러한 결과는 지시학습이 언어적 단서에 대한 반응성을 향상시키는 대신, 과제 전반에 걸친 행동 안정성을 저해할 수 있음을 시사한다. 더 나아가, LLM의 의사결정 행동은 과제 구조와 언어적 표현에 의해 공동으로 형성된다는 점을 보여주며, 고위험 환경이나 의사결정 지원 맥락에 LLM을 적용하기에 앞서 체계적인 행동 평가가 필요함을 강조한다.

    더보기

    목차 (Table of Contents)

    • 1. Introduction 1
    • 2. Related Work 5
    • 2.1. LLMs in Decision-Making and High-Stakes Applications 5
    • 2.2. Risky Decision-Making Tasks Applied to LLMs 6
    • 1. Introduction 1
    • 2. Related Work 5
    • 2.1. LLMs in Decision-Making and High-Stakes Applications 5
    • 2.2. Risky Decision-Making Tasks Applied to LLMs 6
    • 2.3. Sensitivity to Prompt Formulation and Linguistic Framing 8
    • 2.4. Data Contamination Concerns in LLM Evaluation 9
    • 3. Methodological Framework for Risk Preference Evaluation 11
    • 3.1. Risky Decision-Making Tasks 11
    • 3.1.1. Balloon Analogue Risk Task (BART) 13
    • 3.1.2. Columbia Card Task (CCT) 17
    • 3.1.3. Game of Dice Task (GDT) 21
    • 3.2. Linguistic Framing 25
    • 3.2.1. Linguistic Framing in Risky Decision-Making: Prior Approaches and Limitations 25
    • 3.2.2. Design Principles for Linguistic Framing 26
    • 3.2.3. Implementation and Examples of Linguistic Framing 29
    • 3.3. Evaluation Setup 31
    • 3.3.1. Models 31
    • 3.3.2. Experimental Settings 34
    • 4. Validation Study: Assessing Data Contamination 35
    • 4.1. Rationale for Contamination Checks 35
    • 4.2. Experimental Design for Contamination Checks 37
    • 4.3. Results of RQ0 40
    • 4.4. Contamination Risk: Evaluation and Experimental Safeguards 43
    • 5. Empirical Analysis of Risk Behavior in LLMs 45
    • 5.1. RQ1 — Cross-Task Consistency 47
    • 5.1.1. Partial Consistency under Gain Framing 49
    • 5.1.2. Reversal Patterns under Loss Framing 49
    • 5.1.3. Lack of Structure under Neutral Framing 50
    • 5.1.4. Summary of RQ1 50
    • 5.2. RQ2 — Framing Sensitivity 51
    • 5.2.1. Model-Level Framing Patterns 52
    • 5.2.2. Quantifying Framing Effects: FEI and FSI 54
    • 5.2.3. Summary of RQ2 58
    • 5.3. Instruction Tuning and Risk Behavior 59
    • 5.3.1. Impact of Instruction Tuning on Cross-Task Consistency 62
    • 5.3.2. Impact of Instruction Tuning on Framing Sensitivity 66
    • 5.3.3. Summary of Instruction Tuning Effects 75
    • 6. Conclusion 77
    • Bibliography 81
    • Appendix 89
    • Abstract in Korean 95
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼