최근 대량 문자 서비스를 통한 도박, 불법 대출, 주식 리딩방 등의 상업적 목적의 스팸 문자가 증가하면서 이로 인한 피해가 지속되고 있다. 이러한 문제를 해결하고자 다양한 방법들이 연구...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
https://www.riss.kr/link?id=T17075574
서울 : 고려대학교 SW·AI융합대학원, 2024
학위논문(석사) -- 고려대학교 SW·AI융합대학원 , 소프트웨어보안학과 , 2024. 8
2024
한국어
스팸문자 ; 딥러닝 ; 연합학습 ; On-Device AI
서울
x,47p ; 26 cm
지도교수: 유헌창
I804:11009-000000289278
0
상세조회0
다운로드최근 대량 문자 서비스를 통한 도박, 불법 대출, 주식 리딩방 등의 상업적 목적의 스팸 문자가 증가하면서 이로 인한 피해가 지속되고 있다. 이러한 문제를 해결하고자 다양한 방법들이 연구...
최근 대량 문자 서비스를 통한 도박, 불법 대출, 주식 리딩방 등의 상업적 목적의 스팸 문자가 증가하면서 이로 인한 피해가 지속되고 있다. 이러한 문제를 해결하고자 다양한 방법들이 연구되고 정책이 마련되어 왔지만, 이를 회피하는 방법 역시 발전하고 있다.
기존의 스팸 필터링 방식들은 주로 사전에 정의된 스팸 키워드를 기반으로 필터링하거나, 자주 등장하는 단어의 출현 빈도수로 스팸 문자를 검출한다. 그러나 이러한 방법들은 해당 주요 키워드를 의도적으로 변형하는 경우 스팸 문자로 필터링 하지 못한다는 한계가 있다. 사전 정의된 스팸 문자 발신번호를 기반으로 해당 발신번호를 차단하는 방식도 있지만, 이는 최근 증가하고 있는 전화번호 변작기를 통해 발신번호를 조작하는 경우에는 효과적으로 대응하지 못한다. 따라서 기존 방식들을 회피하는 콘텐츠까지도 효율적으로 분석할 방안이 요구되고 있다.
이러한 요구에 부응하여 본 연구에서는 모바일 환경에서 스팸 콘텐츠가 포함된 문자 메시지를 효과적으로 탐지하기 위해 딥러닝, On-Device AI기술, 그리고 연합학습을 결합한 풀-스택(Full-Stack) 시스템을 개발하였다. 본 시스템은 딥러닝 기반의 객체 탐지 및 OCR(Optical Character Recognition, 광학 문자 인식) 기술, 자연어 처리 기술을 결합하여 스팸 메시지를 분류하는 모델을 구현하였으며, 이를 통해 높은 분류 정확도를 달성하였다. 특히 스팸 문자 필터링을 회피하기 위해 의도적으로 변형한 문자 메시지에 대한 분류에도 효과적임을 입증하였다.
또한 On-Device AI 기술을 활용하여 연합학습(Federated Learning)을 모바일 환경에서 실현 가능하게 하였다. 각 사용자의 모바일 기기가 독립적으로 데이터 처리와 학습을 수행함으로써 데이터를 중앙 서버로 직접 공유하지 않아 데이터의 프라이버시 보호를 강화하면서도 중앙 서버의 저장 공간과 트래픽 부담을 감소시킬 수 있는 이점이 있다. 본 연구는 프라이버시 보호 측면의 보안성 강화와 함께 높은 처리 효율성을 갖춘 스팸 메시지 탐지 기술의 새로운 방향을 제시한다.
다국어 초록 (Multilingual Abstract)
Spam texts via bulk texting services are on the rise, and the harm caused by spam texts advertising gambling, illegal loans, stock lending rooms, and more continues to grow. While various methods have been researched and policies have been put in plac...
Spam texts via bulk texting services are on the rise, and the harm caused by spam texts advertising gambling, illegal loans, stock lending rooms, and more continues to grow. While various methods have been researched and policies have been put in place to address these issues, methods to avoid them are also evolving.
Traditional spam filtering methods mainly filter based on predefined spam keywords, or detect spam texts based on the frequency of occurrence of common words. However, these methods are limited in that they fail to filter out intentional variations of these keywords as spam. There are also methods that block spam text callers based on predefined spam text caller IDs, but these methods do not effectively respond to caller ID manipulation through recent phone number modifiers. Therefore, there is a need for an efficient way to analyze even content that evades existing methods.
In this study, we developed a full-stack system that combines deep learning, on-device AI technology, and federated learning to effectively detect text messages containing spam content in the mobile environment. The system combines deep learning-based object detection, optical character recognition (OCR) technology, and natural language processing to implement a model for classifying spam messages, which achieves high classification accuracy. In particular, it proved to be effective in classifying text messages that were intentionally modified to evade spam filtering.
We also utilized on-device AI technology to make federated learning feasible in a mobile environment. By performing data processing and learning independently on each
user's mobile device, the data is not directly shared to the central server, which has the advantage of enhancing data privacy while reducing the storage space and traffic burden on the central server. This research provides a new direction for spam message detection technology with high processing efficiency while enhancing security in terms of privacy protection.
목차 (Table of Contents)