RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Co-optimization of Oxide Semiconductor-based Synaptic Device and Training Algorithms for Robust Analog In-Memory Computing = 강건한 아날로그 인메모리 컴퓨팅을 위한 산화물 반도체 기반 시냅스 소자와 학습 알고리즘의 공동 최적화

    한글로보기

    https://www.riss.kr/link?id=T17449977

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    현대 컴퓨팅 시스템은 메모리와 연산 유닛의 물리적 분리로 인해 발생하는 오랜 von Neumann 병목 문제에 지속적으로 직면해 있다. 이러한 한계는 데이터 집약적인 인공지능(AI) 워크로드가 증가함에 따라 더욱 심각해지고 있다. 아날로그 인메모리 컴퓨팅(AIMC)은 시냅스 크로스바 어레이 내에서 데이터 저장과 연산을 통합함으로써 불필요한 데이터 이동을 제거하고, 이 병목 현상을 완화할 수 있는 유망한 대안으로 부상하였다. 최근 AIMC 연구의 상당 부분은 딥 뉴럴 네트워크(DNN) 추론 가속에 집중되어 왔으나, 수조 개의 파라미터를 갖는 대규모 AI 모델의 등장으로 인해 계산 부담은 점차 학습 단계로 이동하고 있으며, 이 과정에서 발생하는 학습 비용만으로도 수억 달러에 이를 수 있다. 이에 따라, 효율적인 온칩 학습과 적응적 가중치 업데이트가 가능한 AIMC 하드웨어의 개발은 지속 가능한 AI 구현을 위해 필수적인 과제로 대두되고 있다.
    본 학위논문은 정확하고 효율적인 온칩 학습이 가능한 소자–알고리즘 공동 최적화 AIMC 플랫폼을 개발함으로써 이러한 문제를 해결하고자 한다. 디지털 하드웨어에서 가중치 업데이트가 결정론적으로 적용되는 것과 달리, AIMC 하드웨어에서는 펄스 기반 업데이트 방식이 사용되므로, 학습 정확도는 소자 고유의 비이상성에 매우 민감하게 영향을 받는다. 이에 본 연구에서는 AIMC 스택의 여러 계층에 걸쳐 나타나는 한계를 체계적으로 극복하기 위해, 신뢰성 있는 소자 설계, 알고리즘적 보상 기법, 그리고 학습을 고려한 동작 기법을 통합적으로 제시한다.
    제1장에서는 실질적인 AIMC 학습을 지원할 수 있는 시냅스 소자 구조가 무엇인지에 대한 근본적인 질문을 다룬다. 이를 위해, 성숙하고 산업적으로 검증된 공정 기술을 기반으로 제작된 산화물 반도체와 커패시터를 결합한 새로운 6T1C 시냅스 소자를 설계하여 높은 내구성을 확보하였다. 커패시터 기반 시냅스에서 본질적으로 발생하는 장기 리텐션 열화를 해결하기 위해, 소자 특성에 최적화된 알고리즘을 개발하였으며, 이를 통해 신뢰성 있고 정확한 가중치 업데이트를 달성하기 위해서는 소자–알고리즘 공동 최적화가 필수적임을 입증하였다.
    제2장에서는 다수의 병렬 업데이트 펄스가 인가되는 크로스바 어레이 환경에서 아날로그 시냅스 소자에 광범위하게 발생하지만 종종 간과되어 온 쓰기 교란(write disturbance) 문제를 다룬다. 이러한 교란 현상은 명확한 물리적 원리에 기반하여 정량화하기 어렵기 때문에, 학습 정확도에 미치는 영향이 충분히 탐구되지 않았다. 본 장에서는 앞서 개발된 6T1C 소자를 기반으로 쓰기 교란 메커니즘을 체계적으로 분석하고, 6T1C 소자 동작 특성과 rTT 학습 알고리즘에 대한 심층적인 이해를 바탕으로 의도적 펄스 인가, 바이어스 최적화, 그리고 펄스 시퀀스 스케줄링의 세 가지 단순하면서도 효과적인 동작 기법을 제안한다.
    마지막 장에서는 소자 간 변동성이 존재하는 실제 하드웨어 칩 환경에서, 소자 맞춤형 알고리즘과 동작 기법을 포함한 기존 학습 프레임워크가 성공적으로 적용 가능한지를 평가한다. 개별 IGZO TFT가 높은 균일성을 보인다는 통념과 달리, 본 연구는 시냅스 셀 구조와 TFT의 동작 영역에 따라 변동성 민감도가 크게 증가할 수 있음을 보인다. 이를 극복하기 위해 기존 6T1C 구조에 단일 트랜지스터를 추가한 변동성 인지 7T1C 시냅스 소자를 개발하였다. 제안된 소자는 기존 구조의 동작 메커니즘을 유지하면서도 변동성에 대한 강인성을 크게 향상시켜, 이전에 개발된 학습 프레임워크를 수정 없이 적용할 수 있도록 한다.
    종합적으로, 본 학위논문은 확장 가능한 온칩 학습이 가능한 변동성 내성 및 에너지 효율적인 AIMC 하드웨어를 위한 포괄적인 기반을 확립한다. 소자 설계, 알고리즘 개발, 그리고 동작 최적화를 아우르는 통합적 접근을 통해, 차세대 AI 응용을 위한 AIMC 플랫폼의 발전에 필요한 핵심 통찰과 실질적인 방법론을 제시한다.
    번역하기

    현대 컴퓨팅 시스템은 메모리와 연산 유닛의 물리적 분리로 인해 발생하는 오랜 von Neumann 병목 문제에 지속적으로 직면해 있다. 이러한 한계는 데이터 집약적인 인공지능(AI) 워크로드가 증...

    현대 컴퓨팅 시스템은 메모리와 연산 유닛의 물리적 분리로 인해 발생하는 오랜 von Neumann 병목 문제에 지속적으로 직면해 있다. 이러한 한계는 데이터 집약적인 인공지능(AI) 워크로드가 증가함에 따라 더욱 심각해지고 있다. 아날로그 인메모리 컴퓨팅(AIMC)은 시냅스 크로스바 어레이 내에서 데이터 저장과 연산을 통합함으로써 불필요한 데이터 이동을 제거하고, 이 병목 현상을 완화할 수 있는 유망한 대안으로 부상하였다. 최근 AIMC 연구의 상당 부분은 딥 뉴럴 네트워크(DNN) 추론 가속에 집중되어 왔으나, 수조 개의 파라미터를 갖는 대규모 AI 모델의 등장으로 인해 계산 부담은 점차 학습 단계로 이동하고 있으며, 이 과정에서 발생하는 학습 비용만으로도 수억 달러에 이를 수 있다. 이에 따라, 효율적인 온칩 학습과 적응적 가중치 업데이트가 가능한 AIMC 하드웨어의 개발은 지속 가능한 AI 구현을 위해 필수적인 과제로 대두되고 있다.
    본 학위논문은 정확하고 효율적인 온칩 학습이 가능한 소자–알고리즘 공동 최적화 AIMC 플랫폼을 개발함으로써 이러한 문제를 해결하고자 한다. 디지털 하드웨어에서 가중치 업데이트가 결정론적으로 적용되는 것과 달리, AIMC 하드웨어에서는 펄스 기반 업데이트 방식이 사용되므로, 학습 정확도는 소자 고유의 비이상성에 매우 민감하게 영향을 받는다. 이에 본 연구에서는 AIMC 스택의 여러 계층에 걸쳐 나타나는 한계를 체계적으로 극복하기 위해, 신뢰성 있는 소자 설계, 알고리즘적 보상 기법, 그리고 학습을 고려한 동작 기법을 통합적으로 제시한다.
    제1장에서는 실질적인 AIMC 학습을 지원할 수 있는 시냅스 소자 구조가 무엇인지에 대한 근본적인 질문을 다룬다. 이를 위해, 성숙하고 산업적으로 검증된 공정 기술을 기반으로 제작된 산화물 반도체와 커패시터를 결합한 새로운 6T1C 시냅스 소자를 설계하여 높은 내구성을 확보하였다. 커패시터 기반 시냅스에서 본질적으로 발생하는 장기 리텐션 열화를 해결하기 위해, 소자 특성에 최적화된 알고리즘을 개발하였으며, 이를 통해 신뢰성 있고 정확한 가중치 업데이트를 달성하기 위해서는 소자–알고리즘 공동 최적화가 필수적임을 입증하였다.
    제2장에서는 다수의 병렬 업데이트 펄스가 인가되는 크로스바 어레이 환경에서 아날로그 시냅스 소자에 광범위하게 발생하지만 종종 간과되어 온 쓰기 교란(write disturbance) 문제를 다룬다. 이러한 교란 현상은 명확한 물리적 원리에 기반하여 정량화하기 어렵기 때문에, 학습 정확도에 미치는 영향이 충분히 탐구되지 않았다. 본 장에서는 앞서 개발된 6T1C 소자를 기반으로 쓰기 교란 메커니즘을 체계적으로 분석하고, 6T1C 소자 동작 특성과 rTT 학습 알고리즘에 대한 심층적인 이해를 바탕으로 의도적 펄스 인가, 바이어스 최적화, 그리고 펄스 시퀀스 스케줄링의 세 가지 단순하면서도 효과적인 동작 기법을 제안한다.
    마지막 장에서는 소자 간 변동성이 존재하는 실제 하드웨어 칩 환경에서, 소자 맞춤형 알고리즘과 동작 기법을 포함한 기존 학습 프레임워크가 성공적으로 적용 가능한지를 평가한다. 개별 IGZO TFT가 높은 균일성을 보인다는 통념과 달리, 본 연구는 시냅스 셀 구조와 TFT의 동작 영역에 따라 변동성 민감도가 크게 증가할 수 있음을 보인다. 이를 극복하기 위해 기존 6T1C 구조에 단일 트랜지스터를 추가한 변동성 인지 7T1C 시냅스 소자를 개발하였다. 제안된 소자는 기존 구조의 동작 메커니즘을 유지하면서도 변동성에 대한 강인성을 크게 향상시켜, 이전에 개발된 학습 프레임워크를 수정 없이 적용할 수 있도록 한다.
    종합적으로, 본 학위논문은 확장 가능한 온칩 학습이 가능한 변동성 내성 및 에너지 효율적인 AIMC 하드웨어를 위한 포괄적인 기반을 확립한다. 소자 설계, 알고리즘 개발, 그리고 동작 최적화를 아우르는 통합적 접근을 통해, 차세대 AI 응용을 위한 AIMC 플랫폼의 발전에 필요한 핵심 통찰과 실질적인 방법론을 제시한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Modern computing systems continue to face the long-standing von Neumann bottleneck stemming from the physical separation between memory and processing units. This limitation becomes increasingly critical for today’s data-intensive artificial intelligence (AI) workloads. Analog in-memory computing (AIMC) has emerged as a promising alternative by integrating data storage and computation within synaptic crossbar arrays, thereby eliminating unnecessary data movement and alleviating this bottleneck. While much of the recent progress in AIMC has focused on accelerating deep neural network (DNN) inference, the rise of large-scale AI models with trillions of parameters has shifted computational demands toward the training phase, where training costs alone can exceed hundreds of millions of dollars. Consequently, the development of AIMC hardware capable of efficient on-chip training and adaptive weight updates has become essential for realizing sustainable AI.
    This dissertation addresses these challenges by developing device–algorithm co-optimized AIMC platforms capable of accurate and efficient on-chip learning. Unlike digital hardware, where weight updates are deterministically applied, AIMC hardware relies on pulse-based update schemes, making training accuracy highly sensitive to intrinsic device non-idealities. My research therefore integrates reliable device design, algorithmic compensation strategies, and training-aware operational schemes to systematically overcome the limitations encountered across multiple layers of the AIMC stack.
    My first chapter investigates the fundamental question of what type of synaptic device architecture can support practical AIMC training. To answer this, a novel 6T1C synaptic device was designed by combining an oxide semiconductor with a capacitor, both fabricated using mature and industry-proven process technologies to ensure high endurance. To address long-term retention degradation inherent to capacitor-based synapses, a device-specific algorithm was developed, demonstrating the critical importance of device–algorithm co-optimization for achieving reliable and accurate weight updates.
    My second chapter focuses on write disturbance, a pervasive yet often overlooked issue in crossbar arrays where analog synaptic devices undergo numerous parallel update pulses. Because these disturbances are challenging to quantify based on clear physical principles, their impact on training accuracy has remained largely unexplored. Building upon the previously developed 6T1C device, this work provides a systematic analysis of write disturbance mechanisms and introduces three simple yet highly effective operational schemes—intentional pulse application, bias optimization, and pulse-sequence scheduling—derived from a detailed understanding of 6T1C device behavior and the rTT training algorithm.
    The final chapter evaluates whether the established training framework—including device-tailored algorithms and operational schemes—can be successfully applied to real hardware chips exhibiting device-to-device variation. Contrary to the assumption that high uniformity in individual IGZO TFTs directly translates to uniformity in synaptic devices, this work shows that variation sensitivity increases significantly with the synaptic cell structure and the operating region of the TFTs. To overcome this limitation, a variation-aware 7T1C synaptic device was developed by adding a single transistor to the original 6T1C architecture. The proposed device preserves the operational mechanism of the conventional structure while substantially enhancing robustness under variation, enabling seamless application of the previously developed training framework without modification.
    Collectively, this dissertation establishes a comprehensive foundation for variation-tolerant, energy-efficient AIMC hardware capable of scalable on-chip training. Through an integrated approach spanning device design, algorithm development, and operational optimization, it provides key insights and practical methodologies for advancing AIMC platforms toward next-generation AI applications.
    번역하기

    Modern computing systems continue to face the long-standing von Neumann bottleneck stemming from the physical separation between memory and processing units. This limitation becomes increasingly critical for today’s data-intensive artificial intelli...

    Modern computing systems continue to face the long-standing von Neumann bottleneck stemming from the physical separation between memory and processing units. This limitation becomes increasingly critical for today’s data-intensive artificial intelligence (AI) workloads. Analog in-memory computing (AIMC) has emerged as a promising alternative by integrating data storage and computation within synaptic crossbar arrays, thereby eliminating unnecessary data movement and alleviating this bottleneck. While much of the recent progress in AIMC has focused on accelerating deep neural network (DNN) inference, the rise of large-scale AI models with trillions of parameters has shifted computational demands toward the training phase, where training costs alone can exceed hundreds of millions of dollars. Consequently, the development of AIMC hardware capable of efficient on-chip training and adaptive weight updates has become essential for realizing sustainable AI.
    This dissertation addresses these challenges by developing device–algorithm co-optimized AIMC platforms capable of accurate and efficient on-chip learning. Unlike digital hardware, where weight updates are deterministically applied, AIMC hardware relies on pulse-based update schemes, making training accuracy highly sensitive to intrinsic device non-idealities. My research therefore integrates reliable device design, algorithmic compensation strategies, and training-aware operational schemes to systematically overcome the limitations encountered across multiple layers of the AIMC stack.
    My first chapter investigates the fundamental question of what type of synaptic device architecture can support practical AIMC training. To answer this, a novel 6T1C synaptic device was designed by combining an oxide semiconductor with a capacitor, both fabricated using mature and industry-proven process technologies to ensure high endurance. To address long-term retention degradation inherent to capacitor-based synapses, a device-specific algorithm was developed, demonstrating the critical importance of device–algorithm co-optimization for achieving reliable and accurate weight updates.
    My second chapter focuses on write disturbance, a pervasive yet often overlooked issue in crossbar arrays where analog synaptic devices undergo numerous parallel update pulses. Because these disturbances are challenging to quantify based on clear physical principles, their impact on training accuracy has remained largely unexplored. Building upon the previously developed 6T1C device, this work provides a systematic analysis of write disturbance mechanisms and introduces three simple yet highly effective operational schemes—intentional pulse application, bias optimization, and pulse-sequence scheduling—derived from a detailed understanding of 6T1C device behavior and the rTT training algorithm.
    The final chapter evaluates whether the established training framework—including device-tailored algorithms and operational schemes—can be successfully applied to real hardware chips exhibiting device-to-device variation. Contrary to the assumption that high uniformity in individual IGZO TFTs directly translates to uniformity in synaptic devices, this work shows that variation sensitivity increases significantly with the synaptic cell structure and the operating region of the TFTs. To overcome this limitation, a variation-aware 7T1C synaptic device was developed by adding a single transistor to the original 6T1C architecture. The proposed device preserves the operational mechanism of the conventional structure while substantially enhancing robustness under variation, enabling seamless application of the previously developed training framework without modification.
    Collectively, this dissertation establishes a comprehensive foundation for variation-tolerant, energy-efficient AIMC hardware capable of scalable on-chip training. Through an integrated approach spanning device design, algorithm development, and operational optimization, it provides key insights and practical methodologies for advancing AIMC platforms toward next-generation AI applications.

    더보기

    목차 (Table of Contents)

    • Chapter 1. Introduction 1
    • 1.1 Study Background 1
    • 1.2 Objective and Thesis overview 3
    • 1.3 Biography 8
    • Chapter 1. Introduction 1
    • 1.1 Study Background 1
    • 1.2 Objective and Thesis overview 3
    • 1.3 Biography 8
    • Chapter 2. Device-Algorithm Co-Optimization for an On-chip Trainable Capacitor-Based Synaptic Device with IGZO TFT and Retention-Centric Tiki-Taka Algorithm 10
    • 2.1 Introduction 10
    • 2.2 Experimental Methods 14
    • 2.3 Results and Discussions 18
    • 2.3.1 Operational Mechanism of the Synaptic Device 18
    • 2.3.2 Various Properties of a Single Device 23
    • 2.3.3 Training demonstration on a crossbar array 27
    • 2.3.4 Device-specific algorithm 31
    • 2.4 Conclusion 37
    • 2.5 Biography 38
    • Chapter 3. Disturbance-Aware On-chip Training with Mitigation Schemes for Massively Parallel Computing in Analog Deep Learning Accelerator 42
    • 3.1 Introduction 42
    • 3.2 Experimental Methods 45
    • 3.3 Results and Discussions
    • 3.3.1 Disturbance Mechanism Analysis and Mitigation Strategies for 6T1C Synaptic Device 48
    • 3.3.2 Suppression of Unintentional Transistor Activation by DNO 60
    • 3.3.3 Additional Approaches for Reducing Disturbance Magnitude 64
    • 3.3.4 Disturbance-aware training simulation 68
    • 3.4 Conclusion 71
    • 3.5 Biography 73
    • Chapter 4. Variation-Aware InGaZnO Thin Film Transistor-Based Synaptic Device and Training Framework for Robust Analog-In-Memory Computing 77
    • 4.1 Introduction 77
    • 4.2 Experimental Methods 82
    • 4.3 Results and Discussions
    • 4.3.1 Analysis of the Disturbance-Mitigating Mechanism under Device-to-Device Variation 85
    • 4.3.2 Analysis of the symmetry point behavior under Device-to-Device Variation 92
    • 4.3.3 Variation-aware 7T1C Synaptic Device 98
    • 4.3.4 Non-ideality-aware training simulation 104
    • 4.4 Conclusion 112
    • 4.5 Biography 114
    • Chapter 5. Conclusion 124
    • Abstract in Korean 127
    • Acknowledgement 130
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼