RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Enhancing deep-learning-as-a-service with optimized architectures : a framework focused on service-level-objectives

    한글로보기

    https://www.riss.kr/link?id=T17091103

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    As artificial intelligence technology advances, the complexity of AI models increases, and the computing resources required by these models continue to grow. This results in cost issues when scaling AI services.
    Therefore, how to efficiently serve AI models has become a critically important problem, leading to research from various perspectives.
    Among various studies, research on model inference optimization applicable in on-premises environments has been gaining attention. Model inference optimization focuses on minimizing the latency of AI model services and maximizing the performance of models deployed on servers. To achieve this, it aims to determine the most efficient settings for various inference server parameters (such as max batch size, max queue delay time, model instances, dynamic inference, etc.). By optimally configuring the model service architecture, it is possible to enhance the performance of model inference services using the same models and hardware.
    However, there are limitations in configuring the optimal service architecture due to the long exploration time required. Moreover, solely utilizing quality metrics as the criteria for judging service architecture may lead to the problem of absence of objective criteria for selecting optimal configurations.
    Therefore, this study proposes a new integrated metric for determining the optimal service architecture and introduces a service architecture search framework to minimize exploration costs.
    In this paper, experiments are conducted in real-world service environments, demonstrating the effectiveness of the proposed integrated metric and framework in improving performance.
    번역하기

    As artificial intelligence technology advances, the complexity of AI models increases, and the computing resources required by these models continue to grow. This results in cost issues when scaling AI services. Therefore, how to efficiently serve AI ...

    As artificial intelligence technology advances, the complexity of AI models increases, and the computing resources required by these models continue to grow. This results in cost issues when scaling AI services.
    Therefore, how to efficiently serve AI models has become a critically important problem, leading to research from various perspectives.
    Among various studies, research on model inference optimization applicable in on-premises environments has been gaining attention. Model inference optimization focuses on minimizing the latency of AI model services and maximizing the performance of models deployed on servers. To achieve this, it aims to determine the most efficient settings for various inference server parameters (such as max batch size, max queue delay time, model instances, dynamic inference, etc.). By optimally configuring the model service architecture, it is possible to enhance the performance of model inference services using the same models and hardware.
    However, there are limitations in configuring the optimal service architecture due to the long exploration time required. Moreover, solely utilizing quality metrics as the criteria for judging service architecture may lead to the problem of absence of objective criteria for selecting optimal configurations.
    Therefore, this study proposes a new integrated metric for determining the optimal service architecture and introduces a service architecture search framework to minimize exploration costs.
    In this paper, experiments are conducted in real-world service environments, demonstrating the effectiveness of the proposed integrated metric and framework in improving performance.

    더보기

    목차 (Table of Contents)

    • 1. Introduction 1
    • 1.1 Background 1
    • 1.2 Research Scope 3
    • 1.3 Research Purpose 6
    • 1.4 Outlook of this Thesis 7
    • 1. Introduction 1
    • 1.1 Background 1
    • 1.2 Research Scope 3
    • 1.3 Research Purpose 6
    • 1.4 Outlook of this Thesis 7
    • 2. Related Research 8
    • 2.1 Background 8
    • 2.2 Literature Review 9
    • 2.3 Limitations of Previous Research 11
    • 3. Proposed Method 14
    • 3.1 Modified Desirability Function 14
    • 3.2 Service Architecture Search Framework 21
    • 4. Experiments 24
    • 4.1 Experimental Setting 24
    • 4.2 Experimental Results 26
    • 5. Conclusion 43
    • 5.1 Conclusion 43
    • 5.2 Limitations and Future Work 46
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼