RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기
    KCI등재

    엔트로피-인식형 동적 극 분할을 통한 LLM 생성물의 워터마킹 탐지 성능 최적화 = Optimizing Watermark Detection Performance in LLM-Generated Content via Entropy-Aware Dynamic Polarity Partitioning

    한글로보기

    https://www.riss.kr/link?id=A109984568

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    대규모 언어 모델(LLM)의 확산에 따라 생성 텍스트의 출처를 판별하는 워터마킹 기술의 중요성이 대두되고 있다. 본 연구는 고정된 분할 비율을 사용하는 기존 BiMarker 방식의 탐지 성능 한계를 개선하기 위해, 통계 기반 극성 분할 비율 설정을 통해 최적 비율을 선택하는 동적 극 분할(DPP) 기법을 제안한다. 제안 기법은 OPT-1.3B 및 StarCoder 모델, C4 및 HumanEval·MBPP 데이터셋을 기반으로 탐지 성능, 삽입 일관성, 토큰 분포 다양성 측면에서 실험을 수행하였다. 특히 DPP 설정 7(γ = 0.65 / 0.50 / 0.35)은 모든 환경에서 우수한 성능을 나타냈으며, 단어 순서 변경 및 동의어 치환 공격에 대해서도 25% 이내의 성능 저하로 높은 내성을 보였다. 본 연구는 DPP의 실용성과 강건성을 실험적으로 입증하였으며, LLM 기반 텍스트의 신뢰성 확보를 위한 효과적인 대안을 제시한다.
    번역하기

    대규모 언어 모델(LLM)의 확산에 따라 생성 텍스트의 출처를 판별하는 워터마킹 기술의 중요성이 대두되고 있다. 본 연구는 고정된 분할 비율을 사용하는 기존 BiMarker 방식의 탐지 성능 한계...

    대규모 언어 모델(LLM)의 확산에 따라 생성 텍스트의 출처를 판별하는 워터마킹 기술의 중요성이 대두되고 있다. 본 연구는 고정된 분할 비율을 사용하는 기존 BiMarker 방식의 탐지 성능 한계를 개선하기 위해, 통계 기반 극성 분할 비율 설정을 통해 최적 비율을 선택하는 동적 극 분할(DPP) 기법을 제안한다. 제안 기법은 OPT-1.3B 및 StarCoder 모델, C4 및 HumanEval·MBPP 데이터셋을 기반으로 탐지 성능, 삽입 일관성, 토큰 분포 다양성 측면에서 실험을 수행하였다. 특히 DPP 설정 7(γ = 0.65 / 0.50 / 0.35)은 모든 환경에서 우수한 성능을 나타냈으며, 단어 순서 변경 및 동의어 치환 공격에 대해서도 25% 이내의 성능 저하로 높은 내성을 보였다. 본 연구는 DPP의 실용성과 강건성을 실험적으로 입증하였으며, LLM 기반 텍스트의 신뢰성 확보를 위한 효과적인 대안을 제시한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    With the rapid proliferation of large language models (LLMs), watermarking techniques for identifying the origin of generated text have become increasingly important. This study proposes a Dynamic Polarity Partitioning (DPP) method that selects the optimal polarity ratio from predefined statistical settings, addressing the detection limitations of the fixed-ratio-based BiMarker approach. The proposed method is evaluated using the OPT-1.3B and StarCoder models with the C4, HumanEval, and MBPP datasets, focusing on detection performance, watermarking consistency, and token distribution diversity. Experimental results show that DPP setting 7 (γ = 0.65 / 0.50 / 0.35) achieves superior performance across all evaluation criteria. Furthermore, under perturbation attacks such as word order shuffling and synonym substitution, the detection performance decreased by no more than 25%, demonstrating strong robustness. These findings validate the practicality and resilience of DPP and present an effective solution for ensuring the reliability of LLM-generated text.
    번역하기

    With the rapid proliferation of large language models (LLMs), watermarking techniques for identifying the origin of generated text have become increasingly important. This study proposes a Dynamic Polarity Partitioning (DPP) method that selects the op...

    With the rapid proliferation of large language models (LLMs), watermarking techniques for identifying the origin of generated text have become increasingly important. This study proposes a Dynamic Polarity Partitioning (DPP) method that selects the optimal polarity ratio from predefined statistical settings, addressing the detection limitations of the fixed-ratio-based BiMarker approach. The proposed method is evaluated using the OPT-1.3B and StarCoder models with the C4, HumanEval, and MBPP datasets, focusing on detection performance, watermarking consistency, and token distribution diversity. Experimental results show that DPP setting 7 (γ = 0.65 / 0.50 / 0.35) achieves superior performance across all evaluation criteria. Furthermore, under perturbation attacks such as word order shuffling and synonym substitution, the detection performance decreased by no more than 25%, demonstrating strong robustness. These findings validate the practicality and resilience of DPP and present an effective solution for ensuring the reliability of LLM-generated text.

    더보기

    동일학술지(권/호) 다른 논문

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼