RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    대규모 가속기 네트워크 설계를 위한 시뮬레이션 모델 및 설계 공간 탐색 = Simulation Model and Design Space Exploration for Large-Scale Accelerator Network Design

    한글로보기

    https://www.riss.kr/link?id=T17380388

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 연구는 대규모 가속기 네트워크에서 각 연산 단위에 파이프라인된 레 이어를 매핑하는 레이어 파이프라이닝 공간적 매핑(LP-SM)을 통해 대규모 언어 모델(LLM)과 같은 대규모 DNN 워크로드를 실행한다. 그러나 설계 공 간이 방대해 설계 공간 탐색(DSE)이 어렵고, 사이클-정확 시뮬레이션의 과 도한 시간 소요가 병목이 된다. 따라서 본 논문은 하드웨어 동작 관찰에 기 반해 패킷 단위 근사(PGA) 모델을 구축하여, 기존 Gemini[12]와 같은 해석 적 모델 대비 4.25배 더 정확한 성능 예측을 제공하면서 사이클-정확 모델 보다 4.32배 빠르게 동작하는 데에 성공한다. 하지만 샘플링 효율이 부족한 경우 시뮬레이션 속도 향상 만으로는 충분한 DSE 결과 성능의 이득을 충분 히 얻지 못하는데, 이를 해결하기 위해 적응형 탐색 방식을 제안한다. 이는 기존의 Simba[4] 혹은 Gemini와 같은 방식으로 먼저 설계 공간을 빠르게 탐 색한 뒤, 다시 원래 공간으로 재확장하는 적응형 탐색을 제안한다. 동일한 실행시간 예산 내에서 적응형 탐색은 최대 13.1% 더 높은 성능의 설계 옵션 을 식별한다. 본 논문에서 제안하는 성능 예측 모델과 탐색 방식을 모두 적 용할 수 있는 프레임워크를 통해 평균적으로 총 17.3% 더 우수한 설계를 찾 을 수 있다.

    Recent work executes large DNN workloads (e.g., LLMs) on large-scale accelerator networks by mapping pipelined layers onto each processing element via layer-pipelined spatial mapping (LP-SM). However, the vast design space makes design space exploration (DSE) challenging, and the wall-clock time of cycle-accurate simulation becomes the bottleneck. We therefore build a packet-grained approximation (PGA) model from observed hardware behavior that provides performance predictions 4.25× more accurate than the analytical model Gemini [12] while running 4.32× faster than a cycle-accurate model. Yet when sampling efficiency is limited, faster simulation alone does not yield commensurate DSE gains. To address this, we propose adaptive exploration; following Simba [4] and Gemini, it first sweeps a constrained subspace and then re-expands to the original space to recover global optima. Under the same time budget, adaptive exploration identifies design options with up to 13.1% higher performance. Using a unified framework that integrates the proposed prediction model and search method, we find designs that are on average 17.3% better.
    번역하기

    최근 연구는 대규모 가속기 네트워크에서 각 연산 단위에 파이프라인된 레 이어를 매핑하는 레이어 파이프라이닝 공간적 매핑(LP-SM)을 통해 대규모 언어 모델(LLM)과 같은 대규모 DNN 워크로드...

    최근 연구는 대규모 가속기 네트워크에서 각 연산 단위에 파이프라인된 레 이어를 매핑하는 레이어 파이프라이닝 공간적 매핑(LP-SM)을 통해 대규모 언어 모델(LLM)과 같은 대규모 DNN 워크로드를 실행한다. 그러나 설계 공 간이 방대해 설계 공간 탐색(DSE)이 어렵고, 사이클-정확 시뮬레이션의 과 도한 시간 소요가 병목이 된다. 따라서 본 논문은 하드웨어 동작 관찰에 기 반해 패킷 단위 근사(PGA) 모델을 구축하여, 기존 Gemini[12]와 같은 해석 적 모델 대비 4.25배 더 정확한 성능 예측을 제공하면서 사이클-정확 모델 보다 4.32배 빠르게 동작하는 데에 성공한다. 하지만 샘플링 효율이 부족한 경우 시뮬레이션 속도 향상 만으로는 충분한 DSE 결과 성능의 이득을 충분 히 얻지 못하는데, 이를 해결하기 위해 적응형 탐색 방식을 제안한다. 이는 기존의 Simba[4] 혹은 Gemini와 같은 방식으로 먼저 설계 공간을 빠르게 탐 색한 뒤, 다시 원래 공간으로 재확장하는 적응형 탐색을 제안한다. 동일한 실행시간 예산 내에서 적응형 탐색은 최대 13.1% 더 높은 성능의 설계 옵션 을 식별한다. 본 논문에서 제안하는 성능 예측 모델과 탐색 방식을 모두 적 용할 수 있는 프레임워크를 통해 평균적으로 총 17.3% 더 우수한 설계를 찾 을 수 있다.

    Recent work executes large DNN workloads (e.g., LLMs) on large-scale accelerator networks by mapping pipelined layers onto each processing element via layer-pipelined spatial mapping (LP-SM). However, the vast design space makes design space exploration (DSE) challenging, and the wall-clock time of cycle-accurate simulation becomes the bottleneck. We therefore build a packet-grained approximation (PGA) model from observed hardware behavior that provides performance predictions 4.25× more accurate than the analytical model Gemini [12] while running 4.32× faster than a cycle-accurate model. Yet when sampling efficiency is limited, faster simulation alone does not yield commensurate DSE gains. To address this, we propose adaptive exploration; following Simba [4] and Gemini, it first sweeps a constrained subspace and then re-expands to the original space to recover global optima. Under the same time budget, adaptive exploration identifies design options with up to 13.1% higher performance. Using a unified framework that integrates the proposed prediction model and search method, we find designs that are on average 17.3% better.

    더보기

    목차 (Table of Contents)

    • 표목차 ⅱ
    • 그림목차 ⅲ
    • 국문초록 ⅳ
    • 제1장 서론 1
    • 표목차 ⅱ
    • 그림목차 ⅲ
    • 국문초록 ⅳ
    • 제1장 서론 1
    • 제2장 배경 및 관련 연구 3
    • 제1절 대규모 가속기 네트워크 3
    • 제2절 LP-SM 설계공간 4
    • 제3절 성능 예측기 8
    • 1. 해석적 모델 8
    • 2. 시뮬레이션 모델 9
    • 제3장 패킷 단위 근사 모델 11
    • 제4장 적응형 탐색 15
    • 제5장 평가 19
    • 제1절 실험 설정 19
    • 1. 대규모 가속기 네트워크 구성 19
    • 2. 워크로드 19
    • 3. DSE 구성 20
    • 4. DSE 실행 환경 20
    • 제2절 정확도 및 탐색 속도 비교 20
    • 제3절 탐색 방법 비교 26
    • 제4절 실험 결과 요약 28
    • 제6장 결론 30
    • 참고문헌 31
    • ABSTRACT 35
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼