RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    경량 파이프라인 컨트롤러를 통한 질의 인식 기반 비용 효율적 RAG = Query-Aware Cost-Efficient RAG via a Lightweight Pipeline Controller

    한글로보기

    https://www.riss.kr/link?id=T17380351

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 연구는 검색 증강 생성(Retrieval-Augmented Generation, RAG) 시스템 에서 모든 질의에 동일한 고비용 파이프라인을 적용함으로써 발생하는 비용 과 지연, 품질 간 비효율을 줄이기 위해 질의 특성에 따라 검색 구성을 선 제적으로 조정하는 경량 파이프라인 컨트롤러를 제안한다. 컨트롤러는 질의 길이, 숫자와 고유명사 등장 여부, BM25와 밀집 검색 상위 결과의 중첩 정도와 점수 분포 같은 저비용 신호로 질의 난이도와 애매성을 추정하고, 그에 따라 청킹 전략과 검색 깊이, HyDE 사용 여부, 크로 스 엔코더 리랭킹 사용 여부를 한 번에 결정한다. 이를 통해 쉬운 질의에는 작은 k와 최소 구성을, 어려운 질의에는 넓은 k와 선택적 강화 구성을 배치 하며, 검색과 생성 사이 인터페이스에서 모델 수정 없이 정책을 조정할 수 있도록 한다. BEIR의 FiQA와 NFCorpus 개발 세트를 대상으로 최소 비용 fast 고정 파이프라인, 최대 강화 strong 고정 파이프라인과 비교한 결과, 제안 컨트롤 러는 두 데이터셋 모두에서 평균 지연과 토큰 사용량이 fast와 strong 사이의 중간 수준에 머무르면서 nDCG와 MRR 등 상위 랭크 중심 지표에서는 두 기준선을 대체로 상회하였다. 특히 FiQA에서는 strong과 유사한 검색 성능을 유지하면서 비용을 절반 이하로 줄였고, NFCorpus에서는 fast에 가까운 비용 으로 strong보다 높은 상위 랭크 품질을 보여, 질의와 코퍼스 특성에 따라 검색 깊이와 보강 모듈을 동적으로 조정하는 경량 컨트롤러가 비용 의식적 RAG 운용의 실질적 대안이 될 수 있음을 시사한다.
    번역하기

    본 연구는 검색 증강 생성(Retrieval-Augmented Generation, RAG) 시스템 에서 모든 질의에 동일한 고비용 파이프라인을 적용함으로써 발생하는 비용 과 지연, 품질 간 비효율을 줄이기 위해 질의 특성...

    본 연구는 검색 증강 생성(Retrieval-Augmented Generation, RAG) 시스템 에서 모든 질의에 동일한 고비용 파이프라인을 적용함으로써 발생하는 비용 과 지연, 품질 간 비효율을 줄이기 위해 질의 특성에 따라 검색 구성을 선 제적으로 조정하는 경량 파이프라인 컨트롤러를 제안한다. 컨트롤러는 질의 길이, 숫자와 고유명사 등장 여부, BM25와 밀집 검색 상위 결과의 중첩 정도와 점수 분포 같은 저비용 신호로 질의 난이도와 애매성을 추정하고, 그에 따라 청킹 전략과 검색 깊이, HyDE 사용 여부, 크로 스 엔코더 리랭킹 사용 여부를 한 번에 결정한다. 이를 통해 쉬운 질의에는 작은 k와 최소 구성을, 어려운 질의에는 넓은 k와 선택적 강화 구성을 배치 하며, 검색과 생성 사이 인터페이스에서 모델 수정 없이 정책을 조정할 수 있도록 한다. BEIR의 FiQA와 NFCorpus 개발 세트를 대상으로 최소 비용 fast 고정 파이프라인, 최대 강화 strong 고정 파이프라인과 비교한 결과, 제안 컨트롤 러는 두 데이터셋 모두에서 평균 지연과 토큰 사용량이 fast와 strong 사이의 중간 수준에 머무르면서 nDCG와 MRR 등 상위 랭크 중심 지표에서는 두 기준선을 대체로 상회하였다. 특히 FiQA에서는 strong과 유사한 검색 성능을 유지하면서 비용을 절반 이하로 줄였고, NFCorpus에서는 fast에 가까운 비용 으로 strong보다 높은 상위 랭크 품질을 보여, 질의와 코퍼스 특성에 따라 검색 깊이와 보강 모듈을 동적으로 조정하는 경량 컨트롤러가 비용 의식적 RAG 운용의 실질적 대안이 될 수 있음을 시사한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This paper proposes a lightweight pipeline controller for Retrieval-Augmented Generation (RAG) systems, aiming to mitigate inefficiencies in cost, latency, and quality that arise when a single high-cost retrieval pipeline is uniformly applied to all queries. The controller adaptively configures the retrieval pipeline in advance based on query characteristics, rather than relying on a fixed “one-size-fits-all” configuration. The controller estimates query difficulty and ambiguity from low-cost signals such as query length, the presence of numerals and proper nouns, the degree of overlap and score distribution between BM25 and dense retrieval top results, and coverage and concentration indicators derived from top-ranked candidates. Based on this difficulty estimate, it jointly determines the chunking strategy and retrieval depth, as well as the use of HyDE-based query expansion and cross-encoder reranking. Simple queries are routed to configurations with small k and minimal modules, while more difficult queries are allocated larger k and selectively activated enhancement modules. The controller is positioned at the interface between retrieval and generation, enabling dynamic policy control without modifying or retraining underlying models. Experiments on the FiQA and NFCorpus development sets from the BEIR benchmark compare the proposed controller against a minimum-cost “fast” fixed pipeline and a maximally enhanced “strong” fixed pipeline. Across both datasets, the controller attains average latency and token usage between those of fast and strong, while generally outperforming both baselines on top-rank–oriented metrics such as nDCG and MRR. In FiQA, it achieves retrieval performance comparable to the strong pipeline while reducing cost to less than half, and in NFCorpus it delivers superior top-rank quality at a cost close to the fast pipeline. These results indicate that a lightweight controller that dynamically adjusts retrieval depth and optional modules according to query and corpus characteristics offers a practical alternative for cost-aware operation of RAG systems.
    번역하기

    This paper proposes a lightweight pipeline controller for Retrieval-Augmented Generation (RAG) systems, aiming to mitigate inefficiencies in cost, latency, and quality that arise when a single high-cost retrieval pipeline is uniformly applied to all q...

    This paper proposes a lightweight pipeline controller for Retrieval-Augmented Generation (RAG) systems, aiming to mitigate inefficiencies in cost, latency, and quality that arise when a single high-cost retrieval pipeline is uniformly applied to all queries. The controller adaptively configures the retrieval pipeline in advance based on query characteristics, rather than relying on a fixed “one-size-fits-all” configuration. The controller estimates query difficulty and ambiguity from low-cost signals such as query length, the presence of numerals and proper nouns, the degree of overlap and score distribution between BM25 and dense retrieval top results, and coverage and concentration indicators derived from top-ranked candidates. Based on this difficulty estimate, it jointly determines the chunking strategy and retrieval depth, as well as the use of HyDE-based query expansion and cross-encoder reranking. Simple queries are routed to configurations with small k and minimal modules, while more difficult queries are allocated larger k and selectively activated enhancement modules. The controller is positioned at the interface between retrieval and generation, enabling dynamic policy control without modifying or retraining underlying models. Experiments on the FiQA and NFCorpus development sets from the BEIR benchmark compare the proposed controller against a minimum-cost “fast” fixed pipeline and a maximally enhanced “strong” fixed pipeline. Across both datasets, the controller attains average latency and token usage between those of fast and strong, while generally outperforming both baselines on top-rank–oriented metrics such as nDCG and MRR. In FiQA, it achieves retrieval performance comparable to the strong pipeline while reducing cost to less than half, and in NFCorpus it delivers superior top-rank quality at a cost close to the fast pipeline. These results indicate that a lightweight controller that dynamically adjusts retrieval depth and optional modules according to query and corpus characteristics offers a practical alternative for cost-aware operation of RAG systems.

    더보기

    목차 (Table of Contents)

    • 표목차ⅳ
    • 그림목차ⅴ
    • 국문초록ⅶ
    • 제1장 서론 1
    • 표목차ⅳ
    • 그림목차ⅴ
    • 국문초록ⅶ
    • 제1장 서론 1
    • 제1절 연구 배경과 필요성 1
    • 제2절 연구 목적과 범위 3
    • 제3절 연구 구성과 기여 5
    • 제2장 이론적 배경 7
    • 제1절 RAG 패러다임과 경량 컨트롤러 7
    • 1. Retrieval-Augmented Generation(RAG) 7
    • 2. 경량 컨트롤러 9
    • 제2절 검색 파이프라인 핵심 모듈 11
    • 제3절 평가 프레임과 벤치마크 13
    • 제3장 연구 방법 15
    • 제1절 데이터 전처리 17
    • 1. 데이터셋 선정 및 특성 분석 17
    • 2. 텍스트 정규화 및 필드 병합 19
    • 3. 이중 청킹 전략 구현 20
    • 4. 데이터 저장 경로 및 구조 24
    • 제2절 실험 환경 구성 26
    • 1. 공통 모델 자원 26
    • 2. 컨트롤러 설정 28
    • 3. 인덱스 빌드 및 저장 구조 31
    • 4. 인덱스 레지스트리 및 별칭 체계 32
    • 5. 하이브리드 검색 메커니즘 34
    • 6. 재현성 확보 방안 35
    • 제3절 검색 파이프라인 구축 37
    • 1. 질의 난이도 점수 기반 1차 검색 38
    • 2. 증거 기반 HyDE 게이트 40
    • 3. 최종 문서 선별 및 리랭킹 조정 42
    • 제4절 성능 평가 방법 45
    • 1. 평가 시나리오 구성 45
    • 2. 검색 성능 지표 46
    • 3. 효율성 측정 지표 47
    • 4. 로그 저장과 재현성 48
    • 5. 결과 해석 기준 49
    • 제4장 연구 결과 51
    • 제1절 실험 설정 및 로그 개요 51
    • 제2절 검색 성능 비교 55
    • 제3절 검색 속도 비교 61
    • 제4절 검색 비용 및 효율 비교 65
    • 제5절 컨트롤러 설정 분포 분석 70
    • 제6절 사례 질의를 통한 검색 결과 비교 73
    • 1. FiQA 사례: 소득이 없는 사업의 비용 공제 질의 73
    • 2. NFCorpus 사례: 프로바이오틱스와 정신 건강 질의 77
    • 제5장 결론 및 논의 81
    • 참고문헌 88
    • ABSTRACT 90
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼