RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재후보

    RAG 기반 시스템의 신뢰성과 Jailbreaking 보안 취약성 분석 = Reliability of RAG Systems and an Analysis of Jailbreaking Security Vulnerabilities

    한글로보기

    https://www.riss.kr/link?id=A110181571

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    본 연구는 RAG(Retrieval-Augmented Generation) 기반 LLM이 다양한 jailbreak 공격에 대해 어떠한 보안 특성을 보이는지를 체계적으로 평가하였다. 총 135회의 실험을 수행한 결과, Universal jailbreak 프롬프트의 성공률은 12%로 일반 LLM 대비 낮아 강한 방어성을 보였다. 그러나 type-specific 공격의 성공률은 50%에 달해 RAG 특유의 구조적 취약성이 확인되었다. 특히 ‘규정 허점 악용’ 공격은 100% 성공률을 보이며 RAG가 문서에 존재하지 않는 정보 처리(absence reasoning)에 취약함을 보여주었다. 또한 Research Pretext 공격은 60%의 성공률을 기록해, 연구·보안 목적을 가장한 요청이 RAG 안전 필터를 우회할 수 있음을 나타냈다. 이 결과는 RAG 시스템이 일반적인 jailbreak 전략에는 비교적 강하지만, 도메인 특화된 공격에는 쉽게 노출될 수 있는 이중적 보안 특성을 갖는다는 점을 시사한다.
    번역하기

    본 연구는 RAG(Retrieval-Augmented Generation) 기반 LLM이 다양한 jailbreak 공격에 대해 어떠한 보안 특성을 보이는지를 체계적으로 평가하였다. 총 135회의 실험을 수행한 결과, Universal jailbreak 프롬프트...

    본 연구는 RAG(Retrieval-Augmented Generation) 기반 LLM이 다양한 jailbreak 공격에 대해 어떠한 보안 특성을 보이는지를 체계적으로 평가하였다. 총 135회의 실험을 수행한 결과, Universal jailbreak 프롬프트의 성공률은 12%로 일반 LLM 대비 낮아 강한 방어성을 보였다. 그러나 type-specific 공격의 성공률은 50%에 달해 RAG 특유의 구조적 취약성이 확인되었다. 특히 ‘규정 허점 악용’ 공격은 100% 성공률을 보이며 RAG가 문서에 존재하지 않는 정보 처리(absence reasoning)에 취약함을 보여주었다. 또한 Research Pretext 공격은 60%의 성공률을 기록해, 연구·보안 목적을 가장한 요청이 RAG 안전 필터를 우회할 수 있음을 나타냈다. 이 결과는 RAG 시스템이 일반적인 jailbreak 전략에는 비교적 강하지만, 도메인 특화된 공격에는 쉽게 노출될 수 있는 이중적 보안 특성을 갖는다는 점을 시사한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    This study presents a systematic evaluation of security vulnerabilities in Retrieval-Augmented Generation (RAG)–based large language models when subjected to various types of jailbreak attacks. Across 135 trials, the RAG system exhibited strong robustness against universal jailbreak prompts, achieving a low success rate of 12%, significantly lower than typical LLM performance under similar attacks. In contrast, type-specific attacks achieved a 50% success rate, revealing structural weaknesses inherent to the RAG architecture. The Loophole Exploitation attack reached a 100% success rate, indicating that RAG models are particularly vulnerable when handling absent or unstated information in documents. Additionally, the Research Pretext strategy achieved a 60% success rate, suggesting that context framed as research or security evaluation can effectively bypass safety filters. These findings highlight a dual security nature in RAG systems—robust against general jailbreak patterns yet highly susceptible to domain-targeted attacks—providing important insights for future reliability assessments of document-grounded LLMs.
    번역하기

    This study presents a systematic evaluation of security vulnerabilities in Retrieval-Augmented Generation (RAG)–based large language models when subjected to various types of jailbreak attacks. Across 135 trials, the RAG system exhibited strong robu...

    This study presents a systematic evaluation of security vulnerabilities in Retrieval-Augmented Generation (RAG)–based large language models when subjected to various types of jailbreak attacks. Across 135 trials, the RAG system exhibited strong robustness against universal jailbreak prompts, achieving a low success rate of 12%, significantly lower than typical LLM performance under similar attacks. In contrast, type-specific attacks achieved a 50% success rate, revealing structural weaknesses inherent to the RAG architecture. The Loophole Exploitation attack reached a 100% success rate, indicating that RAG models are particularly vulnerable when handling absent or unstated information in documents. Additionally, the Research Pretext strategy achieved a 60% success rate, suggesting that context framed as research or security evaluation can effectively bypass safety filters. These findings highlight a dual security nature in RAG systems—robust against general jailbreak patterns yet highly susceptible to domain-targeted attacks—providing important insights for future reliability assessments of document-grounded LLMs.

    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼