RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Massively Parallel In-Memory Computing Using IGZO TFT and Capacitor-Based Memory Devices = IGZO 박막 트랜지스터와 커패시터 기반 메모리 소자를 이용한 메모리 내 완전 병렬 컴퓨팅

    한글로보기

    https://www.riss.kr/link?id=T17314937

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    With the accumulation of vast amounts of data, deep learning has advanced rapidly. Deep learning involves two key processes: the inference process, which employs matrix-matrix multiply-accumulate (MAC) operations, and the training process, which utilizes outer product operations. Performing these processes on von Neumann architecture-based digital systems requires computation and memory access proportional to the matrix size. However, in the field of analog in-memory computing (AiMC), memristor-based crossbar arrays have been proposed as a novel solution. These arrays leverage Kirchhoff's laws to perform MAC operations in a single step and enable outer product operations by stochastically generating pulses and updating the conductance of memristors at the crosspoints in one operation. While the advantages of AiMC have driven extensive research, the non-ideal programming characteristics of analog devices have led to a focus on accelerating inference by mapping pre-trained parameters from the cloud onto analog device arrays. However, such cloud-based training approaches pose challenges related to latency and privacy, making the development of on-chip training methodologies and memory devices essential.

    In the first part of this research, a 6T1C device based on six InGaZnO (IGZO) thin-film transistors (TFTs) and a capacitor was designed to meet the required specifications for on-chip training, and fabricated. After fabrication, measurements were conducted to demonstrate that the 6T1C device met the required specifications for on-chip training. A 16×16 6T1C crossbar array combined with an FPGA successfully demonstrated on-chip training with over 97% accuracy on the MNIST dataset. To validate on-chip training in the 6T1C array, detailed electrical characterizations were conducted, particularly focusing on temporal variations in ADC outputs during the training process. These measurements confirmed the feasibility of dynamic weight updates in the fabricated devices.

    The second part of this study addresses the scalability limitations inherent to the 6T1C architecture, particularly when attempting to increase integration density by reducing the size of capacitor. Such reduction introduces severe retention degradation and heightened susceptibility to disturbance effects. To alleviate these issues, novel operation schemes were devised to suppress the impact of disturbances. To harness the intrinsic non-idealities, retention decay and disturbance, as a form of hardware-embedded regularization, on-chip training architecture separating the inference path and training path was adopted. In addition, the excessive regularization effect induced by disturbance in convolutional neural network (CNN) was analyzed. To mitigate this, pulse scheduling method and selective update algorithm were developed, leading to robust training. The proposed system successfully trained a ResNet18 model on the CIFAR-10 dataset to a software-equivalent accuracy level and achieved 100% classification accuracy for checkerboard pattern recognition in on-chip training demonstration.

    In the final part of this work, an energy-efficient synaptic circuit tailored for on-chip training was proposed and fabricated. The design adopts a charge-sharing-based voltage sensing scheme, which inherently enables intrinsic regularization and facilitates on-chip training. Furthermore, a 1.5-bit ADC was introduced to realize a deterministic weight transfer mechanism, replacing the high-resolution ADC typically required in architectures that separate inference and training paths. This proposed approach significantly reduced the number of update pulses by avoiding unnecessary weight update in NVM array, thereby enhancing energy efficiency and prolonging the endurance of NVM devices. The architecture demonstrated stable and effective training behavior, offering a viable alternative to conventional designs with improved power and reliability characteristics.
    번역하기

    With the accumulation of vast amounts of data, deep learning has advanced rapidly. Deep learning involves two key processes: the inference process, which employs matrix-matrix multiply-accumulate (MAC) operations, and the training process, which utili...

    With the accumulation of vast amounts of data, deep learning has advanced rapidly. Deep learning involves two key processes: the inference process, which employs matrix-matrix multiply-accumulate (MAC) operations, and the training process, which utilizes outer product operations. Performing these processes on von Neumann architecture-based digital systems requires computation and memory access proportional to the matrix size. However, in the field of analog in-memory computing (AiMC), memristor-based crossbar arrays have been proposed as a novel solution. These arrays leverage Kirchhoff's laws to perform MAC operations in a single step and enable outer product operations by stochastically generating pulses and updating the conductance of memristors at the crosspoints in one operation. While the advantages of AiMC have driven extensive research, the non-ideal programming characteristics of analog devices have led to a focus on accelerating inference by mapping pre-trained parameters from the cloud onto analog device arrays. However, such cloud-based training approaches pose challenges related to latency and privacy, making the development of on-chip training methodologies and memory devices essential.

    In the first part of this research, a 6T1C device based on six InGaZnO (IGZO) thin-film transistors (TFTs) and a capacitor was designed to meet the required specifications for on-chip training, and fabricated. After fabrication, measurements were conducted to demonstrate that the 6T1C device met the required specifications for on-chip training. A 16×16 6T1C crossbar array combined with an FPGA successfully demonstrated on-chip training with over 97% accuracy on the MNIST dataset. To validate on-chip training in the 6T1C array, detailed electrical characterizations were conducted, particularly focusing on temporal variations in ADC outputs during the training process. These measurements confirmed the feasibility of dynamic weight updates in the fabricated devices.

    The second part of this study addresses the scalability limitations inherent to the 6T1C architecture, particularly when attempting to increase integration density by reducing the size of capacitor. Such reduction introduces severe retention degradation and heightened susceptibility to disturbance effects. To alleviate these issues, novel operation schemes were devised to suppress the impact of disturbances. To harness the intrinsic non-idealities, retention decay and disturbance, as a form of hardware-embedded regularization, on-chip training architecture separating the inference path and training path was adopted. In addition, the excessive regularization effect induced by disturbance in convolutional neural network (CNN) was analyzed. To mitigate this, pulse scheduling method and selective update algorithm were developed, leading to robust training. The proposed system successfully trained a ResNet18 model on the CIFAR-10 dataset to a software-equivalent accuracy level and achieved 100% classification accuracy for checkerboard pattern recognition in on-chip training demonstration.

    In the final part of this work, an energy-efficient synaptic circuit tailored for on-chip training was proposed and fabricated. The design adopts a charge-sharing-based voltage sensing scheme, which inherently enables intrinsic regularization and facilitates on-chip training. Furthermore, a 1.5-bit ADC was introduced to realize a deterministic weight transfer mechanism, replacing the high-resolution ADC typically required in architectures that separate inference and training paths. This proposed approach significantly reduced the number of update pulses by avoiding unnecessary weight update in NVM array, thereby enhancing energy efficiency and prolonging the endurance of NVM devices. The architecture demonstrated stable and effective training behavior, offering a viable alternative to conventional designs with improved power and reliability characteristics.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 인공 신경망 기반 기계학습 기술은 대규모 데이터의 축적에 따라 비약적인 발전을 이루었으며, 그 연산 과정은 주로 곱셈-누적 연산을 중심으로 한 추론 단계와 외적 기반의 학습 단계로 구성된다. 이러한 연산을 기존의 폰 노이만 구조 기반 디지털 전자회로에서 수행할 경우, 연산량과 메모리 접근량이 행렬의 크기에 따라 선형적으로 증가하여 병목 현상이 발생한다. 이에 대한 대안으로, 메모리 내 연산 기술이 주목받고 있으며, 특히 저항 기반 메모리를 이용한 교차배열 구조는 회로 법칙을 활용하여 단일 연산 주기 내에 곱셈-누적 연산을 수행하고, 확률적으로 생성된 전기 펄스를 이용하여 소자의 학습을 완전 병렬적으로 수행함으로써 외적 연산 기반 가중치 갱신을 실행할 수 있는 특징이 있다.

    본 연구에서는 클라우드 기반 학습 방식에서 발생하는 시간 지연 및 개인정보 보호 문제를 해소하기 위해, 칩 내부에서 직접 학습을 수행할 수 있는 새로운 메모리 구조와 학습 전용 회로 아키텍처를 설계하고 이를 제조 및 검증하였다. 첫 번째 단계에서는 산화물 반도체 기반 박막 트랜지스터 여섯 개와 축전기 하나로 구성된 6T1C 메모리 소자를 설계하였으며, 칩 내부 학습에 요구되는 전기적 특성을 만족하도록 소자 제작 및 검증하였다. 이후, 해당 소자를 16행 16열의 교차배열로 구성하고, FPG를 이용한 제어 시스템과 연동하여 실시간 학습이 가능한 하드웨어 플랫폼을 구현하였다. MNIST 이미지 자료를 이용한 칩 내 학습 실험에서 97% 이상의 분류 정확도를 달성하였으며, 학습 과정 중 아날로그-디지털 변환기의 출력 전압의 시간적 변화를 정밀하게 측정하여 실시간 가중치 갱신 기능이 구현 가능함을 실험적으로 입증하였다.

    두 번째 단계에서는 집적도를 높이기 위해 축전 용량을 줄일 경우 발생하는 기억 유지 성능 저하 및 병렬 쓰기 동작으로 인한 방해 문제를 해결하기 위한 새로운 동작 방식들을 고안하였다. 또한, 기억 소자의 비이상성인 유지 감쇠와 쓰기방해 현상을 하드웨어 차원의 기계학습 정규화 효과로 활용하기 위해, 추론 경로와 학습 경로를 분리하는 학습 아키텍처를 제안하였다. 더불어, 합성곱 신경망 학습에서 쓰기 방해에 의해 유도되는 과도한 정규화 효과를 분석하였고, 이를 완화하기 위해 펄스 생성 및 인가 순서를 제어하는 스케줄링 방식과 선택적으로 갱신을 수행하는 알고리즘을 설계하였다. 이를 통해 CIFAR10 데이터에 대한 학습에서 기존 소프트웨어 학습 방식과 동등한 수준의 정확도를 달성하였으며, 격자무늬 영상에 대한 하드웨어 상의 학습 실험에서도 100%의 분류 정확도를 입증하였다.

    마지막으로, 학습 전용 회로의 에너지 효율을 향상시키기 위해 전하 공유 방식의 전압 감지 구조를 채택한 새로운 시냅스 회로를 제안하고 이를 제작하였다. 제안된 구조는 회로 자체적으로 정규화 효과를 유도함으로써 학습 안정성을 높이고, 면적 및 에너지 비용이 높은 고해상도 변환기를 대신하여 1.5비트 해상도의 저전력 변환기를 적용함으로써 면적 및 에너지 효율을 높일뿐만 아니라 결정론적인 가중치 전달을 구현하였다. 이로써 불필요한 가중치 전달을 효과적으로 억제하고, 추론 전용 회로 내 비휘발성 메모리 소자의 갱신 횟수를 줄여 전력 소모를 낮추는 동시에 소자의 수명도 향상시킬 수 있었다.

    본 연구는 새로운 소자 구조, 회로 설계, 공정, 학습 아키텍처 및 알고리즘을 통합적으로 구성하여, 칩 내부 컴퓨팅 시스템의 기술적 기반을 마련하였으며, 향후 인공지능 향 반도체 기술의 발전을 위한 유의미한 방향성을 제시하였다.
    번역하기

    최근 인공 신경망 기반 기계학습 기술은 대규모 데이터의 축적에 따라 비약적인 발전을 이루었으며, 그 연산 과정은 주로 곱셈-누적 연산을 중심으로 한 추론 단계와 외적 기반의 학습 단계...

    최근 인공 신경망 기반 기계학습 기술은 대규모 데이터의 축적에 따라 비약적인 발전을 이루었으며, 그 연산 과정은 주로 곱셈-누적 연산을 중심으로 한 추론 단계와 외적 기반의 학습 단계로 구성된다. 이러한 연산을 기존의 폰 노이만 구조 기반 디지털 전자회로에서 수행할 경우, 연산량과 메모리 접근량이 행렬의 크기에 따라 선형적으로 증가하여 병목 현상이 발생한다. 이에 대한 대안으로, 메모리 내 연산 기술이 주목받고 있으며, 특히 저항 기반 메모리를 이용한 교차배열 구조는 회로 법칙을 활용하여 단일 연산 주기 내에 곱셈-누적 연산을 수행하고, 확률적으로 생성된 전기 펄스를 이용하여 소자의 학습을 완전 병렬적으로 수행함으로써 외적 연산 기반 가중치 갱신을 실행할 수 있는 특징이 있다.

    본 연구에서는 클라우드 기반 학습 방식에서 발생하는 시간 지연 및 개인정보 보호 문제를 해소하기 위해, 칩 내부에서 직접 학습을 수행할 수 있는 새로운 메모리 구조와 학습 전용 회로 아키텍처를 설계하고 이를 제조 및 검증하였다. 첫 번째 단계에서는 산화물 반도체 기반 박막 트랜지스터 여섯 개와 축전기 하나로 구성된 6T1C 메모리 소자를 설계하였으며, 칩 내부 학습에 요구되는 전기적 특성을 만족하도록 소자 제작 및 검증하였다. 이후, 해당 소자를 16행 16열의 교차배열로 구성하고, FPG를 이용한 제어 시스템과 연동하여 실시간 학습이 가능한 하드웨어 플랫폼을 구현하였다. MNIST 이미지 자료를 이용한 칩 내 학습 실험에서 97% 이상의 분류 정확도를 달성하였으며, 학습 과정 중 아날로그-디지털 변환기의 출력 전압의 시간적 변화를 정밀하게 측정하여 실시간 가중치 갱신 기능이 구현 가능함을 실험적으로 입증하였다.

    두 번째 단계에서는 집적도를 높이기 위해 축전 용량을 줄일 경우 발생하는 기억 유지 성능 저하 및 병렬 쓰기 동작으로 인한 방해 문제를 해결하기 위한 새로운 동작 방식들을 고안하였다. 또한, 기억 소자의 비이상성인 유지 감쇠와 쓰기방해 현상을 하드웨어 차원의 기계학습 정규화 효과로 활용하기 위해, 추론 경로와 학습 경로를 분리하는 학습 아키텍처를 제안하였다. 더불어, 합성곱 신경망 학습에서 쓰기 방해에 의해 유도되는 과도한 정규화 효과를 분석하였고, 이를 완화하기 위해 펄스 생성 및 인가 순서를 제어하는 스케줄링 방식과 선택적으로 갱신을 수행하는 알고리즘을 설계하였다. 이를 통해 CIFAR10 데이터에 대한 학습에서 기존 소프트웨어 학습 방식과 동등한 수준의 정확도를 달성하였으며, 격자무늬 영상에 대한 하드웨어 상의 학습 실험에서도 100%의 분류 정확도를 입증하였다.

    마지막으로, 학습 전용 회로의 에너지 효율을 향상시키기 위해 전하 공유 방식의 전압 감지 구조를 채택한 새로운 시냅스 회로를 제안하고 이를 제작하였다. 제안된 구조는 회로 자체적으로 정규화 효과를 유도함으로써 학습 안정성을 높이고, 면적 및 에너지 비용이 높은 고해상도 변환기를 대신하여 1.5비트 해상도의 저전력 변환기를 적용함으로써 면적 및 에너지 효율을 높일뿐만 아니라 결정론적인 가중치 전달을 구현하였다. 이로써 불필요한 가중치 전달을 효과적으로 억제하고, 추론 전용 회로 내 비휘발성 메모리 소자의 갱신 횟수를 줄여 전력 소모를 낮추는 동시에 소자의 수명도 향상시킬 수 있었다.

    본 연구는 새로운 소자 구조, 회로 설계, 공정, 학습 아키텍처 및 알고리즘을 통합적으로 구성하여, 칩 내부 컴퓨팅 시스템의 기술적 기반을 마련하였으며, 향후 인공지능 향 반도체 기술의 발전을 위한 유의미한 방향성을 제시하였다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Table of Contents iv
    • List of Tables vi
    • Abstract i
    • Table of Contents iv
    • List of Tables vi
    • List of Figures vii
    • List of Abbreviations xxii
    • 1. Introduction 1
    • 1.1. Massively Parallel On-chip Training with IGZO TFT and Capacitor-Based Memory Devices 1
    • 1.2. Objective and Chapter Overview 4
    • 1.3. References 6
    • 2. IGZO TFTs and Capacitor-based Memory Device for Analog in Memory Computing 8
    • 2.1. Introduction 8
    • 2.2. Experimental 10
    • 2.3. Results and Discussions 17
    • 2.4. Conclusion 31
    • 2.5. References 33
    • 3. Scaling-Induced Non-Idealities as Implicit Regularization for On-Chip Training 37
    • 3.1. Introduction 37
    • 3.2. Experimental 41
    • 3.3. Results and Discussions 51
    • 3.4. Conclusion 97
    • 3.5. References 99
    • 4. On-Chip Training System with 1.5-bit ADC and Crossbar Synaptic Arrays of IGZO TFT-Based Update Cells Enabling Intrinsic Regularization 103
    • 4.1. Introduction 103
    • 4.2. Experimental 106
    • 4.3. Results and Discussions 108
    • 4.4. Conclusion 124
    • 4.5. References 126
    • 5. Conclusion 130
    • Curriculum Vitae 133
    • List of publications 138
    • Abstract (in Korean) 142
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼