RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Multi-Echelon Inventory Management under Yield, Lead Time and Demand Uncertainty: An Explainable Deep Reinforcement Learning Approach = 수율, 리드타임 및 수요 불확실성 하의 다단계 재고 관리: 설명 가능한 심층 강화학습 접근법

    한글로보기

    https://www.riss.kr/link?id=T17450667

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Global supply chains are increasingly exposed to volatility arising from geopolitical disruptions, climate shocks, and structural vulnerabilities. The semiconductor industry exemplifies these challenges through stochastic yields, variable lead times, and amplified demand fluctuations. While deep reinforcement learning (DRL) offers a promising approach for adaptive inventory control, limited explainability remains a critical barrier to practical adoption.
    This study develops an explainable DRL-based inventory management system for semiconductor supply chains in a Fabless–Foundry–OSAT (FFO) framework. A Proximal Policy Optimization (PPO) agent is benchmarked against calibrated heuristic policies under deterministic conditions, stochastic environments, and zero-shot disruption scenarios, including capacity reductions, demand volatility, and logistics delays.
    To enhance interpretability, we apply a multi-lens analysis combining time-series diagnostics, decision histograms, and SHAP-based feature attribution. The analysis reveals consistent decision patterns: finished-goods inventory dominates order decisions, the foundry fast-line functions as an emergency mechanism, and the OSAT fast-line remains structurally active to protect downstream service levels.
    Results show that the DRL policy achieves an average cost reduction of approximately 18% under stochastic uncertainty while maintaining comparable service levels, whereas deterministic settings yield similar performance. Robustness tests indicate favorable performance under capacity drop and demand amplification scenarios through event-triggered control, but performance degrades under logistics delay shocks, highlighting structural limitations beyond the training distribution.
    번역하기

    Global supply chains are increasingly exposed to volatility arising from geopolitical disruptions, climate shocks, and structural vulnerabilities. The semiconductor industry exemplifies these challenges through stochastic yields, variable lead times, ...

    Global supply chains are increasingly exposed to volatility arising from geopolitical disruptions, climate shocks, and structural vulnerabilities. The semiconductor industry exemplifies these challenges through stochastic yields, variable lead times, and amplified demand fluctuations. While deep reinforcement learning (DRL) offers a promising approach for adaptive inventory control, limited explainability remains a critical barrier to practical adoption.
    This study develops an explainable DRL-based inventory management system for semiconductor supply chains in a Fabless–Foundry–OSAT (FFO) framework. A Proximal Policy Optimization (PPO) agent is benchmarked against calibrated heuristic policies under deterministic conditions, stochastic environments, and zero-shot disruption scenarios, including capacity reductions, demand volatility, and logistics delays.
    To enhance interpretability, we apply a multi-lens analysis combining time-series diagnostics, decision histograms, and SHAP-based feature attribution. The analysis reveals consistent decision patterns: finished-goods inventory dominates order decisions, the foundry fast-line functions as an emergency mechanism, and the OSAT fast-line remains structurally active to protect downstream service levels.
    Results show that the DRL policy achieves an average cost reduction of approximately 18% under stochastic uncertainty while maintaining comparable service levels, whereas deterministic settings yield similar performance. Robustness tests indicate favorable performance under capacity drop and demand amplification scenarios through event-triggered control, but performance degrades under logistics delay shocks, highlighting structural limitations beyond the training distribution.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    글로벌 공급망은 지정학적 갈등, 기후 충격, 구조적 취약성으로 인해 높은 변동성에 노출되고 있으며, 반도체 산업은 확률적 수율, 가변적 리드타임, 수요 증폭 현상이 결합된 대표적 사례이다. 심층 강화학습(Deep Reinforcement Learning, DRL)은 이러한 환경에서 적응적 재고 관리를 가능하게 하지만, 설명 가능성 부족은 실제 산업 적용의 주요 제약으로 남아 있다.
    본 연구는 팹리스–파운드리–OSAT로 구성된 반도체 공급망(FFO 프레임워크)을 대상으로 설명 가능한 DRL 기반 재고 관리 시스템을 제안한다. Proximal Policy Optimization(PPO) 기반 DRL 정책을 설계하고, 이를 결정론적 환경, 확률적 불확실성 환경, 그리고 용량 감소, 수요 변동성 확대, 물류 지연을 포함한 제로샷 교란 시나리오에서 보정된 휴리스틱 정책과 비교 평가하였다.
    설명 가능성 확보를 위해 시계열 진단, 의사결정 히스토그램, SHAP 기반 특성 기여도 분석을 결합한 다중 분석 프레임워크를 적용하였다. 분석 결과, 완제품 재고가 주문 의사결정의 핵심 신호로 작용하며, 파운드리 패스트 라인은 비상 대응 수단으로 선택적으로 활용되고, OSAT 패스트 라인은 하류 서비스 수준 보호를 위해 구조적으로 지속 가동되는 일관된 행동 패턴이 확인되었다.
    실험 결과, 확률적 불확실성 환경에서 DRL 정책은 서비스 수준을 유지하면서 평균 약 18%의 비용 절감 효과를 달성한 반면, 결정론적 환경에서는 휴리스틱 정책과 유사한 성능을 보였다. 강건성 분석에서는 용량 감소 및 수요 증폭 시나리오에서 우수한 성능을 유지하였으나, 학습 분포를 벗어나는 리드타임 지연 충격 하에서는 성능 저하가 관찰되어 구조적 한계를 확인하였다.
    본 연구는 설명 가능성 분석을 통해 DRL 정책의 작동 원리와 취약 구간을 체계적으로 규명함으로써, 불확실성이 높은 반도체 공급망에서 인공지능 기반 재고 관리 정책의 실질적 도입 가능성을 제시한다.
    번역하기

    글로벌 공급망은 지정학적 갈등, 기후 충격, 구조적 취약성으로 인해 높은 변동성에 노출되고 있으며, 반도체 산업은 확률적 수율, 가변적 리드타임, 수요 증폭 현상이 결합된 대표적 사례이...

    글로벌 공급망은 지정학적 갈등, 기후 충격, 구조적 취약성으로 인해 높은 변동성에 노출되고 있으며, 반도체 산업은 확률적 수율, 가변적 리드타임, 수요 증폭 현상이 결합된 대표적 사례이다. 심층 강화학습(Deep Reinforcement Learning, DRL)은 이러한 환경에서 적응적 재고 관리를 가능하게 하지만, 설명 가능성 부족은 실제 산업 적용의 주요 제약으로 남아 있다.
    본 연구는 팹리스–파운드리–OSAT로 구성된 반도체 공급망(FFO 프레임워크)을 대상으로 설명 가능한 DRL 기반 재고 관리 시스템을 제안한다. Proximal Policy Optimization(PPO) 기반 DRL 정책을 설계하고, 이를 결정론적 환경, 확률적 불확실성 환경, 그리고 용량 감소, 수요 변동성 확대, 물류 지연을 포함한 제로샷 교란 시나리오에서 보정된 휴리스틱 정책과 비교 평가하였다.
    설명 가능성 확보를 위해 시계열 진단, 의사결정 히스토그램, SHAP 기반 특성 기여도 분석을 결합한 다중 분석 프레임워크를 적용하였다. 분석 결과, 완제품 재고가 주문 의사결정의 핵심 신호로 작용하며, 파운드리 패스트 라인은 비상 대응 수단으로 선택적으로 활용되고, OSAT 패스트 라인은 하류 서비스 수준 보호를 위해 구조적으로 지속 가동되는 일관된 행동 패턴이 확인되었다.
    실험 결과, 확률적 불확실성 환경에서 DRL 정책은 서비스 수준을 유지하면서 평균 약 18%의 비용 절감 효과를 달성한 반면, 결정론적 환경에서는 휴리스틱 정책과 유사한 성능을 보였다. 강건성 분석에서는 용량 감소 및 수요 증폭 시나리오에서 우수한 성능을 유지하였으나, 학습 분포를 벗어나는 리드타임 지연 충격 하에서는 성능 저하가 관찰되어 구조적 한계를 확인하였다.
    본 연구는 설명 가능성 분석을 통해 DRL 정책의 작동 원리와 취약 구간을 체계적으로 규명함으로써, 불확실성이 높은 반도체 공급망에서 인공지능 기반 재고 관리 정책의 실질적 도입 가능성을 제시한다.

    더보기

    목차 (Table of Contents)

    • Chapter 1. Introduction 1
    • Chapter 2. Literature Review 7
    • Chapter 3. Model and Simuation Setting 11
    • Chapter 4. Numerical Example: Semiconductor Supply Chain 21
    • Chapter 5. Discussion and Future Direction 70
    • Chapter 1. Introduction 1
    • Chapter 2. Literature Review 7
    • Chapter 3. Model and Simuation Setting 11
    • Chapter 4. Numerical Example: Semiconductor Supply Chain 21
    • Chapter 5. Discussion and Future Direction 70
    • Appendix 78
    • Bibliography 82
    • Abstract in Korean 86
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼