RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Applied reinforcement learning for sequential recommender systems

    한글로보기

    https://www.riss.kr/link?id=T14462455

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recently enormous amounts of data have been accumulated in both the on-line and off-line market. This data is used in stock and distribution management fields, proposing to enhance customer satisfaction. In order to improve customer satisfaction, recommender systems are used to provide companies with information about what users prefer. Recommender systems have been developed to find users’ favorite products by using the users’ purchasing histories, individual profiles and social network services. Companies which provide customers with a recommender service apply recommender systems to increase profit. The existing recommender systems usually don’t consider the profit aspect and they are focused on finding the exact products which customers like. Accordingly, this research proposes a recommender system model that considers both customers’ preferences and companies’ profits. We use ‘Reinforcement Learning’ to earn more profit rather than recommending high price products in long-term perspective. Applied Reinforcement Learning for Sequential Recommender Systems learn by seeking future maximum reward. This research sets up reward to price element and companies can get higher profit. Also, applied Reinforcement learning for sequential recommender systems don’t need to learn from the beginning because the data is accumulated sequentially. Thus it learns and recommend real time, using additional data. The model arranges customers’ product purchase information by date, using a two-year retail store’s transaction data and is trained by applying transition probability and price element based on customers’ purchase sequence. This model forms a reward network based on the learning process and recommends products that have maximum reward. In result of experiment, compared with Markov Decision Process which resembles the test model, we find that it results rather lower accuracy but it is expected earnings more than two times of that. Finally, we propose the recommender systems that focus on increasing profit and maintaining the accuracy of recommendation. If companies apply this recommendation model, they can both earn more profit and enhance customers satisfaction.
    번역하기

    Recently enormous amounts of data have been accumulated in both the on-line and off-line market. This data is used in stock and distribution management fields, proposing to enhance customer satisfaction. In order to improve customer satisfaction, reco...

    Recently enormous amounts of data have been accumulated in both the on-line and off-line market. This data is used in stock and distribution management fields, proposing to enhance customer satisfaction. In order to improve customer satisfaction, recommender systems are used to provide companies with information about what users prefer. Recommender systems have been developed to find users’ favorite products by using the users’ purchasing histories, individual profiles and social network services. Companies which provide customers with a recommender service apply recommender systems to increase profit. The existing recommender systems usually don’t consider the profit aspect and they are focused on finding the exact products which customers like. Accordingly, this research proposes a recommender system model that considers both customers’ preferences and companies’ profits. We use ‘Reinforcement Learning’ to earn more profit rather than recommending high price products in long-term perspective. Applied Reinforcement Learning for Sequential Recommender Systems learn by seeking future maximum reward. This research sets up reward to price element and companies can get higher profit. Also, applied Reinforcement learning for sequential recommender systems don’t need to learn from the beginning because the data is accumulated sequentially. Thus it learns and recommend real time, using additional data. The model arranges customers’ product purchase information by date, using a two-year retail store’s transaction data and is trained by applying transition probability and price element based on customers’ purchase sequence. This model forms a reward network based on the learning process and recommends products that have maximum reward. In result of experiment, compared with Markov Decision Process which resembles the test model, we find that it results rather lower accuracy but it is expected earnings more than two times of that. Finally, we propose the recommender systems that focus on increasing profit and maintaining the accuracy of recommendation. If companies apply this recommendation model, they can both earn more profit and enhance customers satisfaction.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    온라인 오프라인 마켓 모두에서 구매정보를 비롯한 다양한 정보가 축적되고 있다. 이러한 많은 양의 정보를 바탕으로 고객의 만족도를 높이기 위한 방법들이 제시되고, 더 나아가 재고관리, 물류관리 등 다양한 분야에서 사용되고 있다. 기업에서는 사용자가 선호하는 정보를 제공하고, 만족도를 높이기 위하여 추천시스템이 사용되고 있다. 추천시스템은 사용자가 좋아할만한 상품을 사용자의 구매기록, 개인 프로필, 사회관계망 서비스 등을 이용하여 발견하고 제공하는 방법으로 발전되어 왔다. 추천을 제공하는 기업의 관점에서는 기업의 수익을 증대시키기 위하여 추천 시스템을 적용하여 사용하고 있다. 하지만 기존 추천시스템은 수익 요소를 고려하지 않고, 고객이 좋아할만한 상품을 더욱 정확히 찾아내는데 초점이 맞추어져 있었다. 이에 본 연구에서는 고객의 선호, 기업 관점에서의 수익의 두 가지 요소를 고려한 추천 모델을 제공한다. 또한 단순한 높은 가격의 상품을 추천하기 보다는 장기적인 관점에서의 높은 수익을 제공하기 위하여 강화학습 방법을 이용하였다. 강화학습을 이용한 추천은 미래의 보상이 최대가 되는 행동으로 모델을 학습하고, 본 논문에서는 보상을 가격요소로 설정하여 추천을 제공함에 있어서 추천이 기업의 관점에서 기존 추천시스템 보다 높은 수익을 얻을 수 있도록 하였다. 또한 강화학습을 이용한 추천 시스템은 데이터가 축적됨에 따라 다시 처음부터 학습할 필요가 없이 추가로 주어진 데이터를 이용하여 학습이 가능하기 때문에 실시간으로 학습과 추천이 가능하다. 본 연구에서 모델의 학습을 위하여 소매점의 2년간의 거래 데이터를 사용하여 고객의 날짜 별 상품 구매 정보를 정리하고, 고객의 상품 구매 순서를 바탕으로 전이확률과 가격요소를 이용하여 모델을 학습하였다. 모델 학습 과정에서 상황에 따른 행동의 보상네트워크를 만들고, 보상이 최대가 되는 상품을 추천한다.
    실험결과, 본 연구에서의 모델과 가장 비슷한 추천 방법인 Markov Decision Process 기반 추천과 정확도, 기대 수익 요소를 비교한 결과 정확도의 관점에서는 Markov Decision Process 기반 추천 방법보다 약간 낮은 형태를 보이지만 기대 수익이 2배 이상 차이 나는 것을 발견하였다.
    결론적으로 본 연구는 기업의 수익이 증대되는 방향으로의 추천 시스템 모델을 만들고 이를 활용한다면 기업의 수익이 높아질 수 있다는 것을 확인하였다.
    번역하기

    온라인 오프라인 마켓 모두에서 구매정보를 비롯한 다양한 정보가 축적되고 있다. 이러한 많은 양의 정보를 바탕으로 고객의 만족도를 높이기 위한 방법들이 제시되고, 더 나아가 재고관리,...

    온라인 오프라인 마켓 모두에서 구매정보를 비롯한 다양한 정보가 축적되고 있다. 이러한 많은 양의 정보를 바탕으로 고객의 만족도를 높이기 위한 방법들이 제시되고, 더 나아가 재고관리, 물류관리 등 다양한 분야에서 사용되고 있다. 기업에서는 사용자가 선호하는 정보를 제공하고, 만족도를 높이기 위하여 추천시스템이 사용되고 있다. 추천시스템은 사용자가 좋아할만한 상품을 사용자의 구매기록, 개인 프로필, 사회관계망 서비스 등을 이용하여 발견하고 제공하는 방법으로 발전되어 왔다. 추천을 제공하는 기업의 관점에서는 기업의 수익을 증대시키기 위하여 추천 시스템을 적용하여 사용하고 있다. 하지만 기존 추천시스템은 수익 요소를 고려하지 않고, 고객이 좋아할만한 상품을 더욱 정확히 찾아내는데 초점이 맞추어져 있었다. 이에 본 연구에서는 고객의 선호, 기업 관점에서의 수익의 두 가지 요소를 고려한 추천 모델을 제공한다. 또한 단순한 높은 가격의 상품을 추천하기 보다는 장기적인 관점에서의 높은 수익을 제공하기 위하여 강화학습 방법을 이용하였다. 강화학습을 이용한 추천은 미래의 보상이 최대가 되는 행동으로 모델을 학습하고, 본 논문에서는 보상을 가격요소로 설정하여 추천을 제공함에 있어서 추천이 기업의 관점에서 기존 추천시스템 보다 높은 수익을 얻을 수 있도록 하였다. 또한 강화학습을 이용한 추천 시스템은 데이터가 축적됨에 따라 다시 처음부터 학습할 필요가 없이 추가로 주어진 데이터를 이용하여 학습이 가능하기 때문에 실시간으로 학습과 추천이 가능하다. 본 연구에서 모델의 학습을 위하여 소매점의 2년간의 거래 데이터를 사용하여 고객의 날짜 별 상품 구매 정보를 정리하고, 고객의 상품 구매 순서를 바탕으로 전이확률과 가격요소를 이용하여 모델을 학습하였다. 모델 학습 과정에서 상황에 따른 행동의 보상네트워크를 만들고, 보상이 최대가 되는 상품을 추천한다.
    실험결과, 본 연구에서의 모델과 가장 비슷한 추천 방법인 Markov Decision Process 기반 추천과 정확도, 기대 수익 요소를 비교한 결과 정확도의 관점에서는 Markov Decision Process 기반 추천 방법보다 약간 낮은 형태를 보이지만 기대 수익이 2배 이상 차이 나는 것을 발견하였다.
    결론적으로 본 연구는 기업의 수익이 증대되는 방향으로의 추천 시스템 모델을 만들고 이를 활용한다면 기업의 수익이 높아질 수 있다는 것을 확인하였다.

    더보기

    목차 (Table of Contents)

    • Contents i
    • Table Contents iii
    • Figure Contents v
    • Abstract vi
    • 1. Introduction 1
    • Contents i
    • Table Contents iii
    • Figure Contents v
    • Abstract vi
    • 1. Introduction 1
    • 2. Related Work 5
    • 2.1 Reinforcement Learning 6
    • 2.2 Revenue maximization recommender systems 10
    • 2.3 Markov Decision Process based recommender systems 11
    • 3. Applied reinforcement learning for sequential recommender systems 13
    • 3.1 Introduction 14
    • 3.2 Overview 16
    • 3.3 Learning Phase 19
    • 3.4 Recommender Phase 22
    • 3.5 Model Evaluation 23
    • 3.6 An Illustrative Example 25
    • 4. Experiments and Results 28
    • 4.1 Data Description 29
    • 4.2 Experimental Setup 30
    • 4.3 Experimental Results 30
    • 5. Conclusion 36
    • 5.1 Summary 37
    • 5.2 Future Works 39
    • 國文抄錄 40
    • Reference 42
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼