RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    전망 이론 기반 강화학습을 통한 투자자 처분 효과의 계산인지 모델 = A Computational Cognitive Model of the Disposition Effect Through Prospect Theory-Based Reinforcement Learning

    한글로보기

    https://www.riss.kr/link?id=T17451776

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 연구는 강화학습을 활용하여 행동 경제학의 핵심 편향인 투자자 처분 효과(Disposition Effect)를 모사하고 그 발현 기제를 분석하는 계산 인지 모델(Computational Cognitive Model)을 제시한다. 기존의 금융 강화학습 연구가 샤프 지수와 같은 객관적 성과 지표의 최적화에 치중해온 것과 달리, 본 연구는 전망 이론(Prospect Theory)의 가치 함수를 보상 체계로 채택하여 인간의 비합리적 의사결정 과정을 공학적으로 재현하고 실증적으로 검증하였다.
    연구를 위해 강화학습 환경 내에 자산별 손익을 추적하는 '심적 회계 추적기'를 구현하였으며, 에이전트의 행동 공간을 Dirichlet 분포로 제약하여 예산 제약 조건 하에서의 연속적인 포트폴리오 최적화를 가능하게 했다. 모델 구조 측면에서는 MLP와 LSTM의 비교를 통해 처분 효과 발현을 위한 시계열 기억 메커니즘의 필수성을 확인하였으며, ‘참조점 의존성’이 형성되는 인지적 과정을 공학적으로 구현하였다.
    실험 결과, 첫째, 자산별로 심적 회계를 분리하는 '좁은 프레임(Narrow Frame)' 환경이 전체 자산 통합 프레임(Global Frame)보다 처분 효과를 더욱 강력하게 유도함을 보였다. 특히 통합 프레임 환경에서는 프레임의 전환이 인지 편향을 완화하는 전략으로 작용할 수 있음을 확인하였다. 둘째, Narrow Frame 하에서 손실 회피 계수(λ)와 처분 효과 간의 '역 U자형' 관계를 발견하고 실증 연구와의 수치적 정합성을 규명하였다. 이는 인간의 실제 λ 값이 '손실 보유(미련)'와 '위험 자산 투매(공포)' 사이에서 미련이 최대화되는 인지적 임계점임을 계산 모델이 수치적으로 재현한 것이다. 셋째, L2 Regularization을 통한 일반화 과정을 거친 후에도 처분 효과가 견고하게 유지됨을 보임으로써, 해당 편향이 모델의 오류가 아닌 가치 구조에서 기인하는 본질적 최적해임을 방증하였다.
    본 연구는 인지 과학적 통찰을 금융 AI 모델에 통합함으로써 행동 경제학 이론의 계산적 검증 도구를 제공할 뿐만 아니라, 패닉 셀링(Panic Selling)과 같은 시장의 비합리적 현상을 예측하고 관리하기 위한 새로운 기술적 토대를 마련하였다는 데 의의가 있다.

    주요어 : 강화 학습, 처분 효과, 전망 이론, 심적 회계, 좁은 프레이밍, 디리클레 분포
    번역하기

    본 연구는 강화학습을 활용하여 행동 경제학의 핵심 편향인 투자자 처분 효과(Disposition Effect)를 모사하고 그 발현 기제를 분석하는 계산 인지 모델(Computational Cognitive Model)을 제시한다. 기존...

    본 연구는 강화학습을 활용하여 행동 경제학의 핵심 편향인 투자자 처분 효과(Disposition Effect)를 모사하고 그 발현 기제를 분석하는 계산 인지 모델(Computational Cognitive Model)을 제시한다. 기존의 금융 강화학습 연구가 샤프 지수와 같은 객관적 성과 지표의 최적화에 치중해온 것과 달리, 본 연구는 전망 이론(Prospect Theory)의 가치 함수를 보상 체계로 채택하여 인간의 비합리적 의사결정 과정을 공학적으로 재현하고 실증적으로 검증하였다.
    연구를 위해 강화학습 환경 내에 자산별 손익을 추적하는 '심적 회계 추적기'를 구현하였으며, 에이전트의 행동 공간을 Dirichlet 분포로 제약하여 예산 제약 조건 하에서의 연속적인 포트폴리오 최적화를 가능하게 했다. 모델 구조 측면에서는 MLP와 LSTM의 비교를 통해 처분 효과 발현을 위한 시계열 기억 메커니즘의 필수성을 확인하였으며, ‘참조점 의존성’이 형성되는 인지적 과정을 공학적으로 구현하였다.
    실험 결과, 첫째, 자산별로 심적 회계를 분리하는 '좁은 프레임(Narrow Frame)' 환경이 전체 자산 통합 프레임(Global Frame)보다 처분 효과를 더욱 강력하게 유도함을 보였다. 특히 통합 프레임 환경에서는 프레임의 전환이 인지 편향을 완화하는 전략으로 작용할 수 있음을 확인하였다. 둘째, Narrow Frame 하에서 손실 회피 계수(λ)와 처분 효과 간의 '역 U자형' 관계를 발견하고 실증 연구와의 수치적 정합성을 규명하였다. 이는 인간의 실제 λ 값이 '손실 보유(미련)'와 '위험 자산 투매(공포)' 사이에서 미련이 최대화되는 인지적 임계점임을 계산 모델이 수치적으로 재현한 것이다. 셋째, L2 Regularization을 통한 일반화 과정을 거친 후에도 처분 효과가 견고하게 유지됨을 보임으로써, 해당 편향이 모델의 오류가 아닌 가치 구조에서 기인하는 본질적 최적해임을 방증하였다.
    본 연구는 인지 과학적 통찰을 금융 AI 모델에 통합함으로써 행동 경제학 이론의 계산적 검증 도구를 제공할 뿐만 아니라, 패닉 셀링(Panic Selling)과 같은 시장의 비합리적 현상을 예측하고 관리하기 위한 새로운 기술적 토대를 마련하였다는 데 의의가 있다.

    주요어 : 강화 학습, 처분 효과, 전망 이론, 심적 회계, 좁은 프레이밍, 디리클레 분포

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study presents a computational cognitive model that leverages reinforcement learning to simulate and analyze the underlying mechanisms of the disposition effect, a core behavioral bias in behavioral economics. Departing from conventional financial reinforcement learning research that has predominantly focused on optimizing objective performance metrics such as the Sharpe ratio, this study adopts the value function from Prospect Theory as its reward framework to computationally reproduce and empirically validate irrational human decision-making processes.
    To facilitate this investigation, a Mental Account Tracker was implemented within the reinforcement learning environment to monitor asset-level gains and losses, while the agent's action space was constrained using a Dirichlet distribution to enable continuous portfolio optimization under budget constraints. Architecturally, a comparative analysis between MLP and LSTM networks confirmed the necessity of temporal memory mechanisms for manifesting the disposition effect and computationally implemented the cognitive processes underlying the formation of reference dependence.
    The experimental findings revealed three key insights. First, narrow frame environments that segregate mental accounts by individual assets induced the disposition effect more strongly than global frame approaches that integrate all assets. Notably, in global frame environments, frame-switching emerged as a viable strategy for mitigating cognitive biases. Second, an inverted U-shaped relationship was discovered between the loss aversion coefficient (λ) and the disposition effect under narrow framing, and numerical consistency with empirical research was established. This demonstrates that the computational model numerically reproduced the cognitive threshold where human actual λ values maximize attachment (loss retention) between "holding losses (attachment)" and "panic selling of risky assets (fear)." Third, the disposition effect remained robust even after generalization through L2 regularization, substantiating that this bias represents an inherent optimal solution arising from the value structure rather than a modeling artifact.
    This research contributes to the field by integrating cognitive scientific insights into financial AI models, thereby providing a computational validation tool for behavioral economics theories. Furthermore, it establishes a novel technological foundation for predicting and managing irrational market phenomena such as panic selling.

    Keywords : Reinforcement Learning, Disposition Effect, Prospect Theory, Mental Accounting, Narrow Framing, Dirichlet Distribution
    번역하기

    This study presents a computational cognitive model that leverages reinforcement learning to simulate and analyze the underlying mechanisms of the disposition effect, a core behavioral bias in behavioral economics. Departing from conventional financia...

    This study presents a computational cognitive model that leverages reinforcement learning to simulate and analyze the underlying mechanisms of the disposition effect, a core behavioral bias in behavioral economics. Departing from conventional financial reinforcement learning research that has predominantly focused on optimizing objective performance metrics such as the Sharpe ratio, this study adopts the value function from Prospect Theory as its reward framework to computationally reproduce and empirically validate irrational human decision-making processes.
    To facilitate this investigation, a Mental Account Tracker was implemented within the reinforcement learning environment to monitor asset-level gains and losses, while the agent's action space was constrained using a Dirichlet distribution to enable continuous portfolio optimization under budget constraints. Architecturally, a comparative analysis between MLP and LSTM networks confirmed the necessity of temporal memory mechanisms for manifesting the disposition effect and computationally implemented the cognitive processes underlying the formation of reference dependence.
    The experimental findings revealed three key insights. First, narrow frame environments that segregate mental accounts by individual assets induced the disposition effect more strongly than global frame approaches that integrate all assets. Notably, in global frame environments, frame-switching emerged as a viable strategy for mitigating cognitive biases. Second, an inverted U-shaped relationship was discovered between the loss aversion coefficient (λ) and the disposition effect under narrow framing, and numerical consistency with empirical research was established. This demonstrates that the computational model numerically reproduced the cognitive threshold where human actual λ values maximize attachment (loss retention) between "holding losses (attachment)" and "panic selling of risky assets (fear)." Third, the disposition effect remained robust even after generalization through L2 regularization, substantiating that this bias represents an inherent optimal solution arising from the value structure rather than a modeling artifact.
    This research contributes to the field by integrating cognitive scientific insights into financial AI models, thereby providing a computational validation tool for behavioral economics theories. Furthermore, it establishes a novel technological foundation for predicting and managing irrational market phenomena such as panic selling.

    Keywords : Reinforcement Learning, Disposition Effect, Prospect Theory, Mental Accounting, Narrow Framing, Dirichlet Distribution

    더보기

    목차 (Table of Contents)

    • 초 록 1
    • 목 차 2
    • 제 1 장 서 론 6
    • 1.1 연구의 배경 및 필요성 6
    • 1.2 연구의 목적 7
    • 초 록 1
    • 목 차 2
    • 제 1 장 서 론 6
    • 1.1 연구의 배경 및 필요성 6
    • 1.2 연구의 목적 7
    • 1.3 연구의 구성 8
    • 제 2 장 이론적 배경 및 관련 연구 9
    • 2.1 인지과학과 행동경제학적 배경 9
    • 2.1.1 전망 이론 9
    • 2.1.2 처분 효과와 심적 회계 (Mental Accounting) 10
    • 2.2 금융 시장에서의 강화학습 11
    • 2.2.1 강화학습의 적용 11
    • 2.2.2 PPO (Proximal Policy Optimization) 11
    • 2.3 관련 연구 및 본 연구의 차별점 11
    • 2.3.1 행동 재무와 AI의 융합 연구 11
    • 2.3.2 본 연구의 차별점 12
    • 제 3 장 강화학습 모델 13
    • 3.1 데이터 구성 및 환경 설계 13
    • 3.1.1 입력 데이터 및 Feature Engineering 13
    • 3.1.2 데이터 전처리 (Preprocessing) 15
    • 3.1.3 마르코프 결정 과정 (MDP) 정의 16
    • 3.2 Dirichlet 분포 기반 LSTM-PPO 모델 17
    • 3.2.1 모델 Architecture 17
    • 3.2.2 LSTM(Long Short-Term Memory) 기반 공유층 18
    • 3.2.3 Actor Head: Dirichlet 분포를 활용한 행동 결정 19
    • 3.2.4 Critic Head 21
    • 3.3 Dirichlet 분포 기반 MLP-PPO 모델 22
    • 3.3.1 모델 Architecture 22
    • 3.4 학습 최적화 기법 (Optimization Techniques) 23
    • 3.4.1 일반화된 어드밴티지 추정 (GAE) 23
    • 3.4.2 PPO의 핵심: Clipped Surrogate Objective 25
    • 3.4.3 수치적 안정성 및 일반화 26
    • 3.4.4 학습 스케줄링 27
    • 3.5 심적회계 추적기 (Mental Account Tracker) 구현 27
    • 3.5.1 심적회계의 준거점 설정 27
    • 3.5.2. 처분 효과 측정 지표 (PGR/PLR) 28
    • 3.5.3 Prospect Theory 보상함수와 연동 28
    • 제 4 장 처분효과 실험 30
    • 4.1 실험 개요 및 목적 30
    • 4.1.1 보상 프레이밍에 따른 실험 환경 설계 30
    • 4.1.2 비교 분석을 위한 에이전트 모델 구성 31
    • 4.2 실험 설계 및 측정 지표 32
    • 4.2.1 실험 환경 및 데이터 32
    • 4.2.2 보상 함수 설정 33
    • 4.2.3 검증 통계량 34
    • 4.3 실험 1: MLP-Global Frame 35
    • 4.3.1 모델 학습 결과 35
    • 4.3.2 모델 테스트 결과 39
    • 4.3.3. 처분 효과 분석 40
    • 4.3.4 실험결과 해석 41
    • 4.4 실험 2: MLP-Narrow Frame 42
    • 4.4.1 모델 학습 결과 42
    • 4.4.2 모델 테스트 결과 46
    • 4.4.3. 처분 효과 분석 46
    • 4.4.4 실험결과 해석 48
    • 4.5 실험 3: LSTM1-Global Frame 48
    • 4.5.1 모델 학습 결과 49
    • 4.5.2 모델 테스트 결과 53
    • 4.5.3. 처분 효과 분석 53
    • 4.5.4 실험결과 해석 55
    • 4.6 실험 4: LSTM1-Narrow Frame 57
    • 4.6.1 모델 학습 결과 57
    • 4.6.2 모델 테스트 결과 61
    • 4.6.3. 처분 효과 62
    • 4.6.4 실험결과 해석 64
    • 4.7 실험 5: LSTM2-Global Frame 65
    • 4.7.1 모델 학습 결과 66
    • 4.7.2 모델 테스트 결과 70
    • 4.7.3. 처분 효과 70
    • 4.7.4 실험결과 해석 72
    • 4.8 실험 6: LSTM2-Narrow Frame 73
    • 4.8.1 모델 학습 결과 73
    • 4.8.2 모델 테스트 결과 77
    • 4.8.3. 처분 효과 77
    • 4.8.4 실험결과 해석 79
    • 4.9 종합 논의 80
    • 4.9.1 Base Case 분석: LSTM1-Global vs. LSTM1-Narrow 80
    • 4.9.2 모델 간 실험결과 비교 81
    • 4.9.3 기억 메커니즘의 역할: MLP와 LSTM의 비교 85
    • 4.9.4 실증적 연구와의 수치적 정합성:λ=2.5 87
    • 4.9.5 Global Frame에서의 처분 효과 완화 및 인지적 통합 88
    • 4.9.6 Narrow frame에서 L2 Regularization의 효과 89
    • 4.9.7 Global frame에서 L2 Regularization의 효과 90
    • 제 5 장 결론 및 제언 91
    • 5.1 연구 요약 91
    • 5.2 연구의 의의 및 기여 93
    • 5.3 향후 연구 과제 95
    • 5.4 맺음말 96
    • Appendix 1: 실험 결과 데이터 97
    • Appendix 2: 통계검정 방법 100
    • Appendix 3: 통계검정 결과 101
    • Appendix 4: Dirichlet 분포 확률밀도함수 107
    • 참고 문헌 108
    • Abstract 110
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼