RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    심층강화학습을 활용한 일본–부산 운항 선박 경로 계획과 탐험–활용 전략 분석 = Deep Reinforcement Learning–Based Vessel Route Planning and Exploration–Exploitation Strategy Analysis for the Japan–Busan Shipping Route

    한글로보기

    https://www.riss.kr/link?id=T17389065

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    항만 물류 산업의 디지털 전환이 가속화되면서 운영 방식이 데이터 기반의 예측 및 의사결정 시스템으로 변화하고 있으나, 해상 운송 환경의 높은 불확실성은 여전히 효율적인 항만 운영을 저해하는 주요 요인으로 작용하고 있다. 이를 해결하기 위해 본 연구는 선박 자동 식별 장치(AIS) 데이터와 심층 강화학습(Deep Reinforcement Learning)을 결합하여 선박의 이동을 경로 계획(Route Planning) 문제로 해석하고, 강화학습 기반 항로 생성 프레임워크를 구현하였다.
    선박의 운항 과정을 연속적인 의사결정 문제로 정의하여 마르코프 결정 과정(Markov Decision Process, MDP)으로 모델링하였으며, 실제 AIS 데이터로부터 추출된 통항 밀도(Traffic Density) 정보를 보상 함수에 반영함으로써 에이전트가 현실적인 항로 패턴을 자율적으로 학습할 수 있는 시뮬레이션 환경을 구성하였다. 특히 방대한 해상 상태 공간과 희소한 보상 구조를 갖는 환경에서 학습 효율을 비교·분석하기 위해 Epsilon-greedy, Boltzmann, UCB(Upper Confidence Bound), Thompson Sampling 등 다양한 탐험–활용 전략을 동일한 조건에서 적용하였다.
    실험 결과, 불확실성을 확률적으로 추정하여 탐색 과정에 반영하는 UCB와 Thompson Sampling 전략이 단순 무작위 기반 탐험 방식 대비 빠른 수렴 속도와 높은 누적 보상을 보였으며, 안정적으로 정책에 수렴하여 일관된 경로 계획 결과를 생성함을 확인하였다. 이는 강화학습 기반 경로 계획 접근은 단순한 위치 추적이나 과거 궤적 모방을 넘어, 선박이 실제로 선택할 가능성이 높은 잠재적 항로를 도출할 수 있음을 보여준다. 이러한 결과는 향후 도착 예정 시간(ETA) 예측의 정밀도 향상과 적시 입항(Just-in-Time Arrival) 운항 전략 수립을 위한 기초 정보로 활용될 수 있으며, 데이터 기반 스마트 항만 운영을 위한 인공지능 적용 가능성을 시사한다.
    번역하기

    항만 물류 산업의 디지털 전환이 가속화되면서 운영 방식이 데이터 기반의 예측 및 의사결정 시스템으로 변화하고 있으나, 해상 운송 환경의 높은 불확실성은 여전히 효율적인 항만 운영을 ...

    항만 물류 산업의 디지털 전환이 가속화되면서 운영 방식이 데이터 기반의 예측 및 의사결정 시스템으로 변화하고 있으나, 해상 운송 환경의 높은 불확실성은 여전히 효율적인 항만 운영을 저해하는 주요 요인으로 작용하고 있다. 이를 해결하기 위해 본 연구는 선박 자동 식별 장치(AIS) 데이터와 심층 강화학습(Deep Reinforcement Learning)을 결합하여 선박의 이동을 경로 계획(Route Planning) 문제로 해석하고, 강화학습 기반 항로 생성 프레임워크를 구현하였다.
    선박의 운항 과정을 연속적인 의사결정 문제로 정의하여 마르코프 결정 과정(Markov Decision Process, MDP)으로 모델링하였으며, 실제 AIS 데이터로부터 추출된 통항 밀도(Traffic Density) 정보를 보상 함수에 반영함으로써 에이전트가 현실적인 항로 패턴을 자율적으로 학습할 수 있는 시뮬레이션 환경을 구성하였다. 특히 방대한 해상 상태 공간과 희소한 보상 구조를 갖는 환경에서 학습 효율을 비교·분석하기 위해 Epsilon-greedy, Boltzmann, UCB(Upper Confidence Bound), Thompson Sampling 등 다양한 탐험–활용 전략을 동일한 조건에서 적용하였다.
    실험 결과, 불확실성을 확률적으로 추정하여 탐색 과정에 반영하는 UCB와 Thompson Sampling 전략이 단순 무작위 기반 탐험 방식 대비 빠른 수렴 속도와 높은 누적 보상을 보였으며, 안정적으로 정책에 수렴하여 일관된 경로 계획 결과를 생성함을 확인하였다. 이는 강화학습 기반 경로 계획 접근은 단순한 위치 추적이나 과거 궤적 모방을 넘어, 선박이 실제로 선택할 가능성이 높은 잠재적 항로를 도출할 수 있음을 보여준다. 이러한 결과는 향후 도착 예정 시간(ETA) 예측의 정밀도 향상과 적시 입항(Just-in-Time Arrival) 운항 전략 수립을 위한 기초 정보로 활용될 수 있으며, 데이터 기반 스마트 항만 운영을 위한 인공지능 적용 가능성을 시사한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    As the digital transformation of the port logistics industry accelerates, operational paradigms are shifting toward data-driven prediction and decision-support systems. Nevertheless, the high level of uncertainty inherent in maritime transportation environments remains a major obstacle to efficient port operations. To address this issue, this study combines Automatic Identification System (AIS) data with deep reinforcement learning and interprets vessel navigation as a route planning problem, implementing a reinforcement learning–based route generation framework.
    The vessel navigation process is formulated as a sequential decision-making problem and modeled within a Markov Decision Process (MDP) framework. Traffic density information extracted from real AIS data is incorporated into the reward function, enabling the agent to autonomously learn realistic routing patterns within a simulated maritime environment. In particular, to compare learning efficiency under large-scale state spaces and sparse reward conditions, multiple exploration–exploitation strategies—including epsilongreedy, Boltzmann exploration, Upper Confidence Bound (UCB), and Thompson Sampling—are applied under identical experimental settings.
    Experimental results show that UCB and Thompson Sampling strategies, which explicitly incorporate probabilistic uncertainty into the exploration process, achieve faster convergence and higher cumulative rewards than simple random exploration schemes, while stably converging to consistent route planning outcomes. These findings demonstrate that a reinforcement learning–based route planning approach can move beyond simple position tracking or imitation of historical trajectories, enabling the derivation of plausible vessel routes that vessels are likely to select in practice. The results further indicate that the proposed approach can serve as foundational information for improving the accuracy of estimated time of arrival (ETA) prediction and supporting just-in-time arrival strategies, thereby highlighting the potential applicability of artificial intelligence in data-driven smart port operations.
    번역하기

    As the digital transformation of the port logistics industry accelerates, operational paradigms are shifting toward data-driven prediction and decision-support systems. Nevertheless, the high level of uncertainty inherent in maritime transportation en...

    As the digital transformation of the port logistics industry accelerates, operational paradigms are shifting toward data-driven prediction and decision-support systems. Nevertheless, the high level of uncertainty inherent in maritime transportation environments remains a major obstacle to efficient port operations. To address this issue, this study combines Automatic Identification System (AIS) data with deep reinforcement learning and interprets vessel navigation as a route planning problem, implementing a reinforcement learning–based route generation framework.
    The vessel navigation process is formulated as a sequential decision-making problem and modeled within a Markov Decision Process (MDP) framework. Traffic density information extracted from real AIS data is incorporated into the reward function, enabling the agent to autonomously learn realistic routing patterns within a simulated maritime environment. In particular, to compare learning efficiency under large-scale state spaces and sparse reward conditions, multiple exploration–exploitation strategies—including epsilongreedy, Boltzmann exploration, Upper Confidence Bound (UCB), and Thompson Sampling—are applied under identical experimental settings.
    Experimental results show that UCB and Thompson Sampling strategies, which explicitly incorporate probabilistic uncertainty into the exploration process, achieve faster convergence and higher cumulative rewards than simple random exploration schemes, while stably converging to consistent route planning outcomes. These findings demonstrate that a reinforcement learning–based route planning approach can move beyond simple position tracking or imitation of historical trajectories, enabling the derivation of plausible vessel routes that vessels are likely to select in practice. The results further indicate that the proposed approach can serve as foundational information for improving the accuracy of estimated time of arrival (ETA) prediction and supporting just-in-time arrival strategies, thereby highlighting the potential applicability of artificial intelligence in data-driven smart port operations.

    더보기

    목차 (Table of Contents)

    • 1. 서론 1
    • 1.1 연구 배경 및 문제 제기 1
    • 1.2 연구 목적 및 기여 2
    • 1.3 논문의 구성 3
    • 2. 관련 연구 4
    • 1. 서론 1
    • 1.1 연구 배경 및 문제 제기 1
    • 1.2 연구 목적 및 기여 2
    • 1.3 논문의 구성 3
    • 2. 관련 연구 4
    • 2.1 항만 운영 예측 연구의 전반적 동향 관련 연구 4
    • 2.2 AIS 기반 선박 항로 예측 관련 연구 5
    • 2.2.1 AIS 데이터의 활용과 품질 관리 5
    • 2.2.2 딥러닝 기반 경로 생성 및 계획 모델링 5
    • 2.2.3 강화학습 적용 및 탐험 전략 6
    • 2.2.4 기존 연구의 한계 및 시사점 7
    • 3. 이론적 배경 8
    • 3.1 마르코프 결정 과정(Markov Decision Process) 8
    • 3.2 심층 강화학습(Deep Reinforcement Learning) 10
    • 3.2.1 Q-Learning의 한계와 함수 근사 10
    • 3.2.2 DQN의 최적화 원리 및 학습 안정화 기법 10
    • 3.3 탐험-활용 전략(ExplorationExploitation Strategies) 11
    • 4. 문제 정의 및 방법론 15
    • 4.1 문제 정의 15
    • 4.2 방법론 17
    • 4.2.1 AIS데이터의 특성 17
    • 4.2.2 DQN 기반 항로 계획 모델 18
    • 5. 실험 26
    • 5.1 실험 환경 26
    • 5.2 실험 및 결과 해석 28
    • 5.2.1 탐험활용 전략 비교 실험 (Seed7) 28
    • 5.2.2 미관측 환경에서의 탐험 전략별 일반화 성능 비교 33
    • 5.2.3 다중 시드 성능 안정성 분석 34
    • 5.2.4 베이스라인 및 강화학습 기반 방법 간 DTW 성능 비교 36
    • 5.2.5 탐험 전략(UCB,ThompsonSampling) 간 통계적 분석 37
    • 6. 결론 및 향후 연구 40
    • 6.1 연구 요약 및 시사점 40
    • 6.2 한계점 및 향후 연구 방향 41
    • 참고문헌 44
    • 영문 초록 48
    • 감사의 글 50
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼