RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    강화학습 기반 교통 신호 제어에서 보상 함수 설계가 지체시간 및 탄소 배출에 미치는 영향 분석 = Analysis of the Effects of Reward Function Design on Vehicle Delay and Carbon Emissions in Reinforcement Learning Based Traffic Signal Control

    한글로보기

    https://www.riss.kr/link?id=T17557176

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study aims to quantitatively analyze the effects of reward function design on vehicle delay and CO2 emissions in reinforcement learning-based traffic signal control. Previous studies on reinforcement learning-based signal control have mainly focused on improving traffic operational efficiency, such as reducing delay and increasing throughput. However, the effects of signal control strategies on environmental performance, particularly those associated with vehicle stopping, restarting, acceleration, and deceleration behaviors, have not been sufficiently examined. Accordingly, this study develops a DQN-based traffic signal control model that considers both traffic efficiency and environmental sustainability, and evaluates the performance differences resulting from alternative reward function designs.
    To achieve this objective, a closed-loop learning environment was established by integrating the VISSIM microscopic traffic simulator with Python through the COM interface. An isolated four-leg signalized intersection was selected as the analysis target, and three scenarios with different traffic demand levels and lane configurations were developed. Scenario A represented a relatively low-demand condition with a total inflow of 1,600 vph on a one-lane-per-approach intersection, Scenario B represented a higher-demand condition with 1,900 vph under the same lane configuration, and Scenario C represented a high-demand condition with 3,000 vph on a two-lane-per-approach intersection. The DQN agent observed the number of residual vehicles, average vehicle delay, CO2 emissions per second, and the current signal cycle as state variables, and selected one of eleven predefined signal cycle alternatives as its action.
    Three reward function types were designed and compared: CASE 1, which minimizes vehicle delay; CASE 2, which minimizes carbon emissions; and CASE 3, which jointly considers both delay and CO2 emissions. Each reward function was applied to the three traffic scenarios, resulting in a total of nine experimental combinations. CO2 emissions were estimated using the VT-Micro model based on the instantaneous speed and acceleration data of individual vehicles extracted from VISSIM. Model performance was evaluated using average delay per vehicle and average CO2 emissions per vehicle. In addition, both the arithmetic mean and the 95th percentile threshold were analyzed to assess overall performance and control stability under unfavorable traffic conditions.
    The results showed that, across all scenarios and reward function types, both vehicle delay and CO2 emissions exhibited large fluctuations during the early learning stage due to exploration, but gradually stabilized as training progressed. In addition, a positive relationship between average vehicle delay and average CO2 emissions was observed in all scenarios. This indicates that longer delays are closely associated with increased stopping time, more frequent restarts, and intensified acceleration and deceleration behavior, all of which contribute to higher emissions.
    Among the three reward function designs, CASE 2 generally demonstrated the most consistent and effective performance. In Scenario A, CASE 2 achieved the lowest values for all evaluation indicators, including mean delay, delay p95, mean CO2 emissions, and CO2 emissions p95. In Scenario C, CASE 2 also outperformed the other cases across all indicators, demonstrating that a carbon-emission-oriented reward function can improve not only environmental performance but also traffic operational efficiency, even under high-demand conditions. In Scenario B, CASE 2 showed the best average performance in both delay and CO2 emissions, although the lowest p95 values were observed in CASE 1 for delay and CASE 3 for CO2 emissions. This suggests that, under near-capacity traffic conditions, average performance and stability under extreme congestion may not always be optimized by the same reward function.
    Overall, the findings confirm that directly incorporating CO2 emissions into the reward structure can positively affect both environmental outcomes and traffic efficiency by reducing unnecessary stops, restarts, and abrupt speed changes. This study demonstrates the applicability of environmentally oriented reinforcement learning-based signal control and highlights the need to move beyond delay-centered signal optimization toward integrated control strategies that consider both mobility and sustainability. Future research should extend the proposed approach to multi-intersection networks, validate it using real-world traffic detection data, and incorporate diverse vehicle types and powertrain characteristics.
    번역하기

    This study aims to quantitatively analyze the effects of reward function design on vehicle delay and CO2 emissions in reinforcement learning-based traffic signal control. Previous studies on reinforcement learning-based signal control have mainly focu...

    This study aims to quantitatively analyze the effects of reward function design on vehicle delay and CO2 emissions in reinforcement learning-based traffic signal control. Previous studies on reinforcement learning-based signal control have mainly focused on improving traffic operational efficiency, such as reducing delay and increasing throughput. However, the effects of signal control strategies on environmental performance, particularly those associated with vehicle stopping, restarting, acceleration, and deceleration behaviors, have not been sufficiently examined. Accordingly, this study develops a DQN-based traffic signal control model that considers both traffic efficiency and environmental sustainability, and evaluates the performance differences resulting from alternative reward function designs.
    To achieve this objective, a closed-loop learning environment was established by integrating the VISSIM microscopic traffic simulator with Python through the COM interface. An isolated four-leg signalized intersection was selected as the analysis target, and three scenarios with different traffic demand levels and lane configurations were developed. Scenario A represented a relatively low-demand condition with a total inflow of 1,600 vph on a one-lane-per-approach intersection, Scenario B represented a higher-demand condition with 1,900 vph under the same lane configuration, and Scenario C represented a high-demand condition with 3,000 vph on a two-lane-per-approach intersection. The DQN agent observed the number of residual vehicles, average vehicle delay, CO2 emissions per second, and the current signal cycle as state variables, and selected one of eleven predefined signal cycle alternatives as its action.
    Three reward function types were designed and compared: CASE 1, which minimizes vehicle delay; CASE 2, which minimizes carbon emissions; and CASE 3, which jointly considers both delay and CO2 emissions. Each reward function was applied to the three traffic scenarios, resulting in a total of nine experimental combinations. CO2 emissions were estimated using the VT-Micro model based on the instantaneous speed and acceleration data of individual vehicles extracted from VISSIM. Model performance was evaluated using average delay per vehicle and average CO2 emissions per vehicle. In addition, both the arithmetic mean and the 95th percentile threshold were analyzed to assess overall performance and control stability under unfavorable traffic conditions.
    The results showed that, across all scenarios and reward function types, both vehicle delay and CO2 emissions exhibited large fluctuations during the early learning stage due to exploration, but gradually stabilized as training progressed. In addition, a positive relationship between average vehicle delay and average CO2 emissions was observed in all scenarios. This indicates that longer delays are closely associated with increased stopping time, more frequent restarts, and intensified acceleration and deceleration behavior, all of which contribute to higher emissions.
    Among the three reward function designs, CASE 2 generally demonstrated the most consistent and effective performance. In Scenario A, CASE 2 achieved the lowest values for all evaluation indicators, including mean delay, delay p95, mean CO2 emissions, and CO2 emissions p95. In Scenario C, CASE 2 also outperformed the other cases across all indicators, demonstrating that a carbon-emission-oriented reward function can improve not only environmental performance but also traffic operational efficiency, even under high-demand conditions. In Scenario B, CASE 2 showed the best average performance in both delay and CO2 emissions, although the lowest p95 values were observed in CASE 1 for delay and CASE 3 for CO2 emissions. This suggests that, under near-capacity traffic conditions, average performance and stability under extreme congestion may not always be optimized by the same reward function.
    Overall, the findings confirm that directly incorporating CO2 emissions into the reward structure can positively affect both environmental outcomes and traffic efficiency by reducing unnecessary stops, restarts, and abrupt speed changes. This study demonstrates the applicability of environmentally oriented reinforcement learning-based signal control and highlights the need to move beyond delay-centered signal optimization toward integrated control strategies that consider both mobility and sustainability. Future research should extend the proposed approach to multi-intersection networks, validate it using real-world traffic detection data, and incorporate diverse vehicle types and powertrain characteristics.

    더보기

    목차 (Table of Contents)

    • 제 1 장 서 론 1
    • 1.1 연구의 배경 및 목적 1
    • 1.2 연구 동향 2
    • 1.2.1 교차로 에너지 절감 연구 2
    • 1.2.2 교차로 강화학습 적용 연구 3
    • 제 1 장 서 론 1
    • 1.1 연구의 배경 및 목적 1
    • 1.2 연구 동향 2
    • 1.2.1 교차로 에너지 절감 연구 2
    • 1.2.2 교차로 강화학습 적용 연구 3
    • 제 2 장 이론적 배경 6
    • 2.1 머신러닝(Machine Learning) 6
    • 2.1.1 머신러닝의 개요 6
    • 2.1.2 머신러닝의 구분 7
    • 2.2 차량 배출량 산정 방법 10
    • 2.2.1 차량 배출량 산정 개요 10
    • 2.2.2 차량 배출량 산정 모델 12
    • 제 3 장 연구 방법 18
    • 3.1 분석 개요 18
    • 3.2 교통 시뮬레이션 모형 20
    • 3.2.1 교통 시뮬레이션 모형 개요 20
    • 3.2.2 VISSIM 시뮬레이션 모형 21
    • 3.3 심층 강화학습 (Deep Reinforcement Learning) 24
    • 3.3.1 심층 강화학습 개요 24
    • 3.3.2 마르코프 결정 과정(Markov Decision Process) 25
    • 3.3.3 Deep Q-Network(DQN) 알고리즘 30
    • 3.4 학습 시나리오 설정 34
    • 3.4.1 학습 시나리오 구성 34
    • 3.4.2 학습 실험 조건 35
    • 3.5 모델 성능 평가 방법 36
    • 3.5.1 평가 개요 36
    • 3.5.2 평가 지표 산정 37
    • 3.5.3 평균 성능 및 상위 5% 임계값 분석 38
    • 3.5.4 보상 함수별 성능 비교 방법 39
    • 3.5.5 지체시간과 CO2 배출량의 관계 분석 39
    • 제 4 장 분석 결과 40
    • 4.1 학습 수렴도 분석 40
    • 4.1.1 CASE 1: 지체 최소화 41
    • 4.1.2 CASE 2: 탄소 배출 최소화 43
    • 4.1.3 CASE 3: 복합 최적화 45
    • 4.2 지체시간과 CO2 배출량 관계 분석 47
    • 4.2.1 시나리오 A 47
    • 4.2.2 시나리오 B 49
    • 4.2.3 시나리오 C 50
    • 4.3 보상 함수별 종합 성능 비교 52
    • 4.3.1 시나리오 A 성능 비교 53
    • 4.3.2 시나리오 B 성능 비교 54
    • 4.3.3 시나리오 C 성능 비교 56
    • 제 5 장 결 론 58
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼