RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    ProRouter: 비용 효율적인 LLM 협업을 위한 프로젝션 기반 라우터 : SVD를 이용한 라우팅 = ProRouter: A Projection-Based Router for Cost-Efficient LLM Collaboration

    한글로보기

    https://www.riss.kr/link?id=T17369885

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recent advances in large language models (LLMs) have greatly improved performance across natural language tasks, but their substantial resource cost poses serious challenges for practical deployment. Although numerous routing methods have been proposed to reduce this cost, most rely on output scores, estimated difficulty, or heuristic rules, and often require additional fine-tuning of the base models. In this paper, we present ProRouter, a projection-based routing framework that efficiently selects between large and small language models (SLMs) using subspace projections of query embeddings. Given an input query, ProRouter first computes its dense embedding using a pre-trained encoder. For each candidate model, it applies singular value decomposition (SVD) to the embeddings of queries labeled according to whether the model’s predictions were correct or incorrect, constructing low-dimensional subspaces that capture the distinctive patterns of success and failure. At inference time, a new query embedding is projected onto these subspaces to obtain similarity- and projection-based features, including cosine similarity to mean embeddings, top-𝑘 similarity, and projection magnitude. These features are concatenated and fed into a lightweight routing classifier to select the most appropriate model, without retraining or fine-tuning of the underlying LLMs. By routing simpler queries to SLMs and complex ones to LLMs, ProRouter substantially reduces computation and latency while maintaining competitive accuracy. Experimental results show that ProRouter achieves large reductions in LLM usage, demonstrating the effectiveness of subspace-projection features for scalable and cost-efficient multi-model collaboration.
    번역하기

    Recent advances in large language models (LLMs) have greatly improved performance across natural language tasks, but their substantial resource cost poses serious challenges for practical deployment. Although numerous routing methods have been propose...

    Recent advances in large language models (LLMs) have greatly improved performance across natural language tasks, but their substantial resource cost poses serious challenges for practical deployment. Although numerous routing methods have been proposed to reduce this cost, most rely on output scores, estimated difficulty, or heuristic rules, and often require additional fine-tuning of the base models. In this paper, we present ProRouter, a projection-based routing framework that efficiently selects between large and small language models (SLMs) using subspace projections of query embeddings. Given an input query, ProRouter first computes its dense embedding using a pre-trained encoder. For each candidate model, it applies singular value decomposition (SVD) to the embeddings of queries labeled according to whether the model’s predictions were correct or incorrect, constructing low-dimensional subspaces that capture the distinctive patterns of success and failure. At inference time, a new query embedding is projected onto these subspaces to obtain similarity- and projection-based features, including cosine similarity to mean embeddings, top-𝑘 similarity, and projection magnitude. These features are concatenated and fed into a lightweight routing classifier to select the most appropriate model, without retraining or fine-tuning of the underlying LLMs. By routing simpler queries to SLMs and complex ones to LLMs, ProRouter substantially reduces computation and latency while maintaining competitive accuracy. Experimental results show that ProRouter achieves large reductions in LLM usage, demonstrating the effectiveness of subspace-projection features for scalable and cost-efficient multi-model collaboration.

    더보기

    목차 (Table of Contents)

    • 1 Introduction 5
    • 2 Related Work 9
    • 2.0.1 Cost-Efficient LLM Cascades and Ensembling 9
    • 2.0.2 LLM Routing Models and Expert Selection 9
    • 2.0.3 Mixture-of-Experts and Model Merging 10
    • 1 Introduction 5
    • 2 Related Work 9
    • 2.0.1 Cost-Efficient LLM Cascades and Ensembling 9
    • 2.0.2 LLM Routing Models and Expert Selection 9
    • 2.0.3 Mixture-of-Experts and Model Merging 10
    • 3 Proposed Method 11
    • 3.0.1 Query Embedding Extraction 11
    • 3.0.2 Model-Specific Correct/Incorrect Subspace Construction 12
    • 3.0.3 Subspace Projection and Feature Engineering 13
    • 3.0.4 Label Construction: Soft and Strict Targets 16
    • 3.0.5 Dynamic Multi-Task Loss 16
    • 4 Experiments 18
    • 4.0.1 Experimental Setup 21
    • 4.0.2 Evaluation Metrics 21
    • 4.0.3 Results and Analysis 22
    • 5 Conclusion 25
    • A Experiment Settings 26
    • A.1 Ablation study 26
    • A.2 Cost model for IRT-specific comparison 26
    • A.3 Supervised data used to build SVD subspaces and train the router 26
    • A.4 Additional OOD Evaluation Results 26
    • A.5 Feature Construction and Embedding 28
    • A.6 Router Model and Training 28
    • A.7 Compute and Infrastructure 29
    • B Prompt Templates 30
    • B.1 CommonsenseQA (CQA) Prompt 30
    • B.2 OpenbookQA Prompt 30
    • B.3 GSM8K Prompt 31
    • B.4 RACE Prompt 33
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼