최근 인공지능을 사물 인터넷 (IoT)과 결합하여 AIoT로 활용하려는 노력이 많다. 인공지능은 많은 연산이 필요하므로, 이를 가속하기 위해 뉴럴 프로세싱 유닛 (NPU)을 사용한다. 하지만 IoT 환경...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
최근 인공지능을 사물 인터넷 (IoT)과 결합하여 AIoT로 활용하려는 노력이 많다. 인공지능은 많은 연산이 필요하므로, 이를 가속하기 위해 뉴럴 프로세싱 유닛 (NPU)을 사용한다. 하지만 IoT 환경...
최근 인공지능을 사물 인터넷 (IoT)과 결합하여 AIoT로 활용하려는 노력이 많다. 인공지능은 많은 연산이 필요하므로, 이를 가속하기 위해 뉴럴 프로세싱 유닛 (NPU)을 사용한다. 하지만 IoT 환경에서 파워 소모와 면적에 대한 제약이 크고, 이를 해결하기 위해 큰 용량의 온 칩 메모리가 요구된다. SRAM은 면적의 제약으로 온 칩 메모리 용량을 무한정 늘릴 수 없지만, 최근 등장한 하이브리드 IGZO/Si을 사용하여 수십 배의 온 칩 메모리 용량이 가능하다. 이렇게 급격히 증가한 온 칩 메모리 용량을 어떻게 활용하는 것이 가장 효율적인지에 관한 연구가 필요하다.
본 연구에서 기존에 온 칩 메모리를 활용하는 데이터 플로우로 제시된 데이터 재사용 (Data Reuse)과 레이어 융합 (Layer Fusion)의 성능을 뉴럴 프로세싱 유닛 시뮬레이터인 SCALE-Sim으로 확인했다. 더 나아가 하이브리드 IGZO/Si을 더욱 효과적으로 활용할 수 있는 회전 버퍼 (Rotation Buffer)를 적용한 웨이트 고정 (Weight-Fixed) 데이터 플로우를 제안했다. 합성곱 신경망, 인코더 기반 트랜스포머, 디코더 기반 트랜스포머에서 데이터 재사용, 레이어 융합, 웨이트 고정이 뉴럴 프로세싱 유닛의 성능에 미치는 영향을 자세히 분석했고, 최종적으로 에너지 효율이 최대 3.09 배 그리고 스루풋이 최대 3.05 배 향상된다는 것을 보였다.
다국어 초록 (Multilingual Abstract)
Recent efforts have focused on integrating AI with the IoT, creating AIoT applications. Since AI requires significant computational resources, NPU is employed to accelerate these processes. However, the power consumption and area constraints in IoT en...
Recent efforts have focused on integrating AI with the IoT, creating AIoT applications. Since AI requires significant computational resources, NPU is employed to accelerate these processes. However, the power consumption and area constraints in IoT environments necessitate the use of large on-chip memory. While the capacity of SRAM-based on-chip memory cannot be indefinitely increased due to area limitations, the recent emergence of Hybrid IGZO/Si technologies enables a substantial increase in on-chip memory capacity, potentially by several orders of magnitude. Therefore, it is crucial to investigate the most
efficient way to utilize the rapidly expanded on-chip memory.
This study investigates the performance of traditional on-chip memory data flows, namely Data Reuse and Layer Fusion, using the NPU simulator SCALE-Sim. Additionally, to further optimize the use of Hybrid IGZO/Si, we propose the Weight-Fixed data flow by incorporating a Rotation Buffer. We analyze the impact of Data Reuse, Layer Fusion, and Weight-Fixed data flows on the performance of NPU in CNNs, encoder-based transformers, and decoder-based transformers. The results demonstrate that these approaches improve energy efficiency by up to 3.09 times and throughput by up to 3.05 times.
목차 (Table of Contents)