최근 연구는 대규모 가속기 네트워크에서 각 연산 단위에 파이프라인된 레 이어를 매핑하는 레이어 파이프라이닝 공간적 매핑(LP-SM)을 통해 대규모 언어 모델(LLM)과 같은 대규모 DNN 워크로드...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
최근 연구는 대규모 가속기 네트워크에서 각 연산 단위에 파이프라인된 레 이어를 매핑하는 레이어 파이프라이닝 공간적 매핑(LP-SM)을 통해 대규모 언어 모델(LLM)과 같은 대규모 DNN 워크로드...
최근 연구는 대규모 가속기 네트워크에서 각 연산 단위에 파이프라인된 레 이어를 매핑하는 레이어 파이프라이닝 공간적 매핑(LP-SM)을 통해 대규모 언어 모델(LLM)과 같은 대규모 DNN 워크로드를 실행한다. 그러나 설계 공 간이 방대해 설계 공간 탐색(DSE)이 어렵고, 사이클-정확 시뮬레이션의 과 도한 시간 소요가 병목이 된다. 따라서 본 논문은 하드웨어 동작 관찰에 기 반해 패킷 단위 근사(PGA) 모델을 구축하여, 기존 Gemini[12]와 같은 해석 적 모델 대비 4.25배 더 정확한 성능 예측을 제공하면서 사이클-정확 모델 보다 4.32배 빠르게 동작하는 데에 성공한다. 하지만 샘플링 효율이 부족한 경우 시뮬레이션 속도 향상 만으로는 충분한 DSE 결과 성능의 이득을 충분 히 얻지 못하는데, 이를 해결하기 위해 적응형 탐색 방식을 제안한다. 이는 기존의 Simba[4] 혹은 Gemini와 같은 방식으로 먼저 설계 공간을 빠르게 탐 색한 뒤, 다시 원래 공간으로 재확장하는 적응형 탐색을 제안한다. 동일한 실행시간 예산 내에서 적응형 탐색은 최대 13.1% 더 높은 성능의 설계 옵션 을 식별한다. 본 논문에서 제안하는 성능 예측 모델과 탐색 방식을 모두 적 용할 수 있는 프레임워크를 통해 평균적으로 총 17.3% 더 우수한 설계를 찾 을 수 있다.
Recent work executes large DNN workloads (e.g., LLMs) on large-scale accelerator networks by mapping pipelined layers onto each processing element via layer-pipelined spatial mapping (LP-SM). However, the vast design space makes design space exploration (DSE) challenging, and the wall-clock time of cycle-accurate simulation becomes the bottleneck. We therefore build a packet-grained approximation (PGA) model from observed hardware behavior that provides performance predictions 4.25× more accurate than the analytical model Gemini [12] while running 4.32× faster than a cycle-accurate model. Yet when sampling efficiency is limited, faster simulation alone does not yield commensurate DSE gains. To address this, we propose adaptive exploration; following Simba [4] and Gemini, it first sweeps a constrained subspace and then re-expands to the original space to recover global optima. Under the same time budget, adaptive exploration identifies design options with up to 13.1% higher performance. Using a unified framework that integrates the proposed prediction model and search method, we find designs that are on average 17.3% better.
목차 (Table of Contents)