RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Model-based Reinforcement Learning for Fatigue-aware Recommendation = 피로도 인식 추천을 위한 모델 기반 강화학습

    한글로보기

    https://www.riss.kr/link?id=T17450630

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    We investigate efficient recommendation strategies that explicitly account for marketing fatigue-the increasing risk of churn induced by repeated exposure to similar items. While a rich literature on bandits and reinforcement learning exists for recommendation, few studies address the long-term impact of recommendations on marketing fatigue and churn. To bridge this gap, we formulate the problem as a Fatigue-aware episodic Markov Decision Process (FA--MDP), incorporating an absorbing churn state, where transition probabilities are defined by a history-dependent logistic function. Building on this formulation, we propose an efficient model-based reinforcement learning algorithm tailored for this environment.

    We first demonstrate that an Upper Confidence Bound-based algorithm offers theoretical regret guarantees in this setting, but it suffers from computational intractability due to an exponentially growing state space.

    To overcome the scalability issue, we introduce an algorithm leveraging Monte Carlo Tree Search (MCTS) to approximate the optimistic policy: Epistemic MCTS for Fatigue-aware MDP (EMCTS-FA). EMCTS-FA employs a theoretically motivated optimistic reward function that incorporates the learned model's epistemic uncertainty to effectively balance the exploration-exploitation trade-off.

    Experiments on both synthetic and real-world data-based environments demonstrate the superior performance of EMCTS-FA over baselines, validating the benefits of epistemic uncertainty-guided exploration in maximizing long-term returns.
    번역하기

    We investigate efficient recommendation strategies that explicitly account for marketing fatigue-the increasing risk of churn induced by repeated exposure to similar items. While a rich literature on bandits and reinforcement learning exists for recom...

    We investigate efficient recommendation strategies that explicitly account for marketing fatigue-the increasing risk of churn induced by repeated exposure to similar items. While a rich literature on bandits and reinforcement learning exists for recommendation, few studies address the long-term impact of recommendations on marketing fatigue and churn. To bridge this gap, we formulate the problem as a Fatigue-aware episodic Markov Decision Process (FA--MDP), incorporating an absorbing churn state, where transition probabilities are defined by a history-dependent logistic function. Building on this formulation, we propose an efficient model-based reinforcement learning algorithm tailored for this environment.

    We first demonstrate that an Upper Confidence Bound-based algorithm offers theoretical regret guarantees in this setting, but it suffers from computational intractability due to an exponentially growing state space.

    To overcome the scalability issue, we introduce an algorithm leveraging Monte Carlo Tree Search (MCTS) to approximate the optimistic policy: Epistemic MCTS for Fatigue-aware MDP (EMCTS-FA). EMCTS-FA employs a theoretically motivated optimistic reward function that incorporates the learned model's epistemic uncertainty to effectively balance the exploration-exploitation trade-off.

    Experiments on both synthetic and real-world data-based environments demonstrate the superior performance of EMCTS-FA over baselines, validating the benefits of epistemic uncertainty-guided exploration in maximizing long-term returns.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    본 논문은 유사한 아이템의 반복적인 노출로 인해 사용자 이탈 위험이 증가하는 현상인 마케팅 피로도를 명시적으로 고려하는 효율적인 추천 전략을 연구한다. 밴딧 및 강화 학습을 추천 시스템에 적용하려는 선행 연구는 많았지만, 추천이 사용자의 피로도와 이탈에 미치는 장기적인 영향을 다룬 연구는 드물다. 이러한 한계를 극복하기 위해, 우리는 피로도를 고려한 마르코프 결정 과정(Fatigue-aware Markov Decision Process, FA-MDP)을 제안한다. 이 모델은 전이 확률이 최근 추천 이력을 기반으로 한 로지스틱 함수에 의해 결정되며, 이탈 상태를 도입하여 자연스럽게 피로도로 인한 이탈 위험을 반영한다.

    먼저, 우리는 Upper Confidence Bound 기반 알고리즘이 이 환경에서 이론적인 후회(regret) 상한을 보장함을 증명하였으나, 상태 공간이 기하급수적으로 증가함에 따라 실질적으로 계산 및 실행이 불가능해지는 문제가 있음을 확인하였다.

    이러한 확장성 문제를 극복하기 위해, 우리는 몬테카를로 트리 탐색을 활용하여 효율적인 탐색의 핵심인 낙관적 정책을 근사하도록 하는 EMCTS-FA (Epistemic Monte Carlo Tree Search for FA-MDP) 알고리즘을 제안한다. EMCTS-FA는 이론으로부터 영감을 받은 낙관적 보상 함수를 사용하여 학습된 모델의 인식론적 불확실성을 반영한 효과적인 탐색을 유도한다.

    인공 및 실제 데이터 기반의 시뮬레이터를 통한 다양한 크기의 환경에서의 실험 결과, EMCTS-FA가 불확실성을 고려하지 않는 알고리즘들에 비해 일관적으로 우수한 성능을 보였으며, 이는 피로도를 관리하고 장기적인 보상을 극대화하는 데 있어 본 연구에서 제안한 인식론적 불확실성에 기반한 탐색이 유효함을 입증한다.
    번역하기

    본 논문은 유사한 아이템의 반복적인 노출로 인해 사용자 이탈 위험이 증가하는 현상인 마케팅 피로도를 명시적으로 고려하는 효율적인 추천 전략을 연구한다. 밴딧 및 강화 학습을 추천 시...

    본 논문은 유사한 아이템의 반복적인 노출로 인해 사용자 이탈 위험이 증가하는 현상인 마케팅 피로도를 명시적으로 고려하는 효율적인 추천 전략을 연구한다. 밴딧 및 강화 학습을 추천 시스템에 적용하려는 선행 연구는 많았지만, 추천이 사용자의 피로도와 이탈에 미치는 장기적인 영향을 다룬 연구는 드물다. 이러한 한계를 극복하기 위해, 우리는 피로도를 고려한 마르코프 결정 과정(Fatigue-aware Markov Decision Process, FA-MDP)을 제안한다. 이 모델은 전이 확률이 최근 추천 이력을 기반으로 한 로지스틱 함수에 의해 결정되며, 이탈 상태를 도입하여 자연스럽게 피로도로 인한 이탈 위험을 반영한다.

    먼저, 우리는 Upper Confidence Bound 기반 알고리즘이 이 환경에서 이론적인 후회(regret) 상한을 보장함을 증명하였으나, 상태 공간이 기하급수적으로 증가함에 따라 실질적으로 계산 및 실행이 불가능해지는 문제가 있음을 확인하였다.

    이러한 확장성 문제를 극복하기 위해, 우리는 몬테카를로 트리 탐색을 활용하여 효율적인 탐색의 핵심인 낙관적 정책을 근사하도록 하는 EMCTS-FA (Epistemic Monte Carlo Tree Search for FA-MDP) 알고리즘을 제안한다. EMCTS-FA는 이론으로부터 영감을 받은 낙관적 보상 함수를 사용하여 학습된 모델의 인식론적 불확실성을 반영한 효과적인 탐색을 유도한다.

    인공 및 실제 데이터 기반의 시뮬레이터를 통한 다양한 크기의 환경에서의 실험 결과, EMCTS-FA가 불확실성을 고려하지 않는 알고리즘들에 비해 일관적으로 우수한 성능을 보였으며, 이는 피로도를 관리하고 장기적인 보상을 극대화하는 데 있어 본 연구에서 제안한 인식론적 불확실성에 기반한 탐색이 유효함을 입증한다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • Chapter 1. Introduction 1
    • 1.1. Related Work 4
    • Chapter 2. Preliminaries 7
    • Abstract i
    • Contents ii
    • Chapter 1. Introduction 1
    • 1.1. Related Work 4
    • Chapter 2. Preliminaries 7
    • 2.1. Notations 7
    • 2.2. Episodic MDP and Regret 7
    • 2.3. Monte Carlo Tree Search 9
    • Chapter 3. Problem Formulation 12
    • Chapter 4. Model-based Algorithm using Dynamic Programming 15
    • 4.1. Upper Confidence Bound-based Algorithm in Fatigue-aware MDP 15
    • 4.2. Estimation of Reward and Transition parameters 16
    • 4.3. Upper Confidence Bound of Value Function 17
    • 4.4. Regret Bound for UCRL-FA 19
    • Chapter 5. Scalable Algorithm with Monte Carlo Tree Search 21
    • 5.1. MCTS for Deep Exploration 21
    • 5.2. Epistemic MCTS for Fatigue-aware MDP (EMCTS-FA) 22
    • Chapter 6. Experiments 25
    • 6.1. Small Environment Experiments 25
    • 6.2. Large Environment Experiments 26
    • 6.3. Results 27
    • Chapter 7. Conclusion 31
    • Appendix 32
    • Bibliography 45
    • Abstract (In Korean) 52
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼