RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Ultra-dynamic Fine-grained Power and Clock Management Techniques for Microprocessors and Machine Learning Accelerators.

    한글로보기

    https://www.riss.kr/link?id=T15667259

    • 저자
    • 발행사항

      Ann Arbor : ProQuest Dissertations & Theses, 2019

    • 학위수여대학

      Northwestern University Electrical and Computer Engineering

    • 수여연도

      2019

    • 작성언어

      영어

    • 주제어
    • 학위

      Ph.D.

    • 페이지수

      143 p.

    • 지도교수/심사위원

      Advisor: Gu, Jie.

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    The increasing prominence of ultra-low power computing applications, such as IoT or wearable devices, require a renewed focus on system-level energy management. Dynamic voltage and frequency scaling (DVFS) has become a widely utilized power management scheme to scale down the microprocessor supply voltage and operation frequency to eliminate the variation protection guardband and save energy consumption. In the commercial SoC chips, each core of the CPU or GPU processors can be powered by independent voltage regulator and clocked by separate clock domain to support flexible power optimization. Although the prior schemes achieve significant energy saving, the operation speed of previous DVFS is still bounded at program level. In another word, the longest exercised critical path delay within the whole program determines the minimum voltage that can be scaled to for each power domain. In fact, there is dynamic timing slack existing within each instruction whenever the longest critical path is not exercised by the runtime instructions. Exploiting the runtime dynamic timing slack can bring promising additional performance/energy gain beyond the conventional DVFS.In this thesis, the ultra-dynamic fine-grained clock and power management scheme operating at instruction level is explored, which lead to significant performance improvement or power saving benefit beyond the conventional program level DVFS management. The instruction-driven adaptive clock management schemes have been explored and developed for (1) a low-power CPU pipeline architecture, (2) a general-purpose GPU architecture, and (3) a PE array based machine learning accelerator architecture. An instruction-driven adaptive clock management through the use of instruction timing encoding and multi-phase all-digital PLL is implemented on a 55nm ARMv7 ISA low power microprocessor. The measurement result shows a 15% performance improvement from proposed clock scheme, which equivalently to up to 28% energy saving benefit. Similar adaptive clocking management scheme is developed on a deep pipelined GPU architecture on a 65nm test chip with hierarchical instruction-driven clocking scheme. The measurement result shows a 18% performance improvement or 30% energy saving from benchmark programs. The dynamic timing slack has also been exploited on the machine learning accelerators by our developed clock chain techniques, which provides multiple clock domains for each PE row and each clock domain loosely synchronized with their neighbors.In addition, a series of fully integrated buck regulator topologies have been developed with novel resonant switching scheme to integrate all the discrete components into silicon, with achieving fast response for the dynamic voltage management. The state-of-the-art power efficiency has been achieved for low power range, with utilizing the fastest switching frequency, i.e. up to 2GHz, reported so far.With the cross-layer design methodology considering from architecture level down to circuit level, the instruction-based dynamic timing slack can be well exploited during runtime. Significant performance improvement or power saving benefits have been obtained on silicon test chips beyond conventional DVFS approaches by our developed ultra-dynamic instruction-driven clock and power management techniques.
    번역하기

    The increasing prominence of ultra-low power computing applications, such as IoT or wearable devices, require a renewed focus on system-level energy management. Dynamic voltage and frequency scaling (DVFS) has become a widely utilized power managemen...

    The increasing prominence of ultra-low power computing applications, such as IoT or wearable devices, require a renewed focus on system-level energy management. Dynamic voltage and frequency scaling (DVFS) has become a widely utilized power management scheme to scale down the microprocessor supply voltage and operation frequency to eliminate the variation protection guardband and save energy consumption. In the commercial SoC chips, each core of the CPU or GPU processors can be powered by independent voltage regulator and clocked by separate clock domain to support flexible power optimization. Although the prior schemes achieve significant energy saving, the operation speed of previous DVFS is still bounded at program level. In another word, the longest exercised critical path delay within the whole program determines the minimum voltage that can be scaled to for each power domain. In fact, there is dynamic timing slack existing within each instruction whenever the longest critical path is not exercised by the runtime instructions. Exploiting the runtime dynamic timing slack can bring promising additional performance/energy gain beyond the conventional DVFS.In this thesis, the ultra-dynamic fine-grained clock and power management scheme operating at instruction level is explored, which lead to significant performance improvement or power saving benefit beyond the conventional program level DVFS management. The instruction-driven adaptive clock management schemes have been explored and developed for (1) a low-power CPU pipeline architecture, (2) a general-purpose GPU architecture, and (3) a PE array based machine learning accelerator architecture. An instruction-driven adaptive clock management through the use of instruction timing encoding and multi-phase all-digital PLL is implemented on a 55nm ARMv7 ISA low power microprocessor. The measurement result shows a 15% performance improvement from proposed clock scheme, which equivalently to up to 28% energy saving benefit. Similar adaptive clocking management scheme is developed on a deep pipelined GPU architecture on a 65nm test chip with hierarchical instruction-driven clocking scheme. The measurement result shows a 18% performance improvement or 30% energy saving from benchmark programs. The dynamic timing slack has also been exploited on the machine learning accelerators by our developed clock chain techniques, which provides multiple clock domains for each PE row and each clock domain loosely synchronized with their neighbors.In addition, a series of fully integrated buck regulator topologies have been developed with novel resonant switching scheme to integrate all the discrete components into silicon, with achieving fast response for the dynamic voltage management. The state-of-the-art power efficiency has been achieved for low power range, with utilizing the fastest switching frequency, i.e. up to 2GHz, reported so far.With the cross-layer design methodology considering from architecture level down to circuit level, the instruction-based dynamic timing slack can be well exploited during runtime. Significant performance improvement or power saving benefits have been obtained on silicon test chips beyond conventional DVFS approaches by our developed ultra-dynamic instruction-driven clock and power management techniques.

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼