RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    검색결과 좁혀 보기

    선택해제
    • 좁혀본 항목 보기순서

      • 원문유무
      • 음성지원유무
      • 학위유형
      • 주제분류
        펼치기
      • 수여기관
        펼치기
      • 발행연도
        펼치기
      • 작성언어
        펼치기
      • 지도교수
        펼치기

    오늘 본 자료

    • 오늘 본 자료가 없습니다.
    더보기
    • 트랜스포머 기반 분류 작업을 위한 시계열 표현 학습

      서재진 인하대학교 대학원 2022 국내석사

      RANK : 2943

      시계열 데이터란 일정한 시간 동안 수집된 일련의 순차적으로 정해진 데이터 셋의 집합을 의미하며 예측, 분류, 이상치 탐지 등에 활용되고 있다. 기존의 시계열 분야의 인공지능 모델에는 RNN(Recurrent Neural Network)을 주로 활용하여 분석을 진행했지만, 최근 Transformer 모델의 개발로 인하여 연구 추세가 변화하고 있다. Transformer 모델은 시계열 데이터 예측에는 좋은 성능을 보이지만, 분류 쪽에서는 상대적으로 부족한 성능을 보인다. 본 논문에서는 시계열 분류를 위한 Transformer 모델에 CLS 토큰을 추가하여 성능 향상에 초점을 맞추었다. 본 논문에서 제안하는 방식은 1) 입력 데이터의 임베딩 방법, 2) 사전 학습 방법이다. 1) 입력 데이터의 임베딩 방법은 총 2가지 방법을 이용한다. 첫 번째는 입력 데이터를 standard scaler를 활용하여 각기 다른 진폭을 가지는 시계열 데이터들을 정규화하여 진폭을 균일하게 만들고 time window 방식으로 데이터의 차원을 변경한 뒤 GRU(Gated Recurrent Unit)를 통하여 Transformer에 입력 토큰으로 활용한다. 두 번째는 GASF(Gramian Angular Summation Field)를 활용하여 입력 데이터를 이미지로 만든 뒤 사전 학습된 컴퓨터 비전 모델을 활용하여 얻어낸 벡터를 Transformer의 CLS 토큰 입력으로 활용한다. 사전 학습 방식은 자연어 분야에서 사용하는 MLM(Masked Language Modeling)과 유사한 방식을 활용한다. 시계열 데이터는 자연어와 다르게 연속 변수로 이루어져 있어서 목적함수 계산 시 MSE(Mean Squared Error)를 활용한다. 입력 토큰의 마스킹 작업 시에 CLS 토큰은 제외하고 나머지 입력 토큰 중 30%를 마스킹하고 마스킹 된 값을 출력단에서 맞추는 형식으로 학습된다. 본 논문에서는 UCR 데이터 셋을 활용하여 총 12개의 서로 다른 모델들과 제안하는 모델의 성능을 비교한다. 제안하는 모델은 85개의 데이터에 대한 평균 정확도 평가에서 최소 1.4% 최대 21.1%까지 성능 향상을 보였다. Time series data refer to a sequentially determined data set collected for a certain period of time and is used for prediction, classification, and outlier detection. Although the existing artificial intelligence models in the field of time series are mainly based on the RNN (Recurrent Neural Network), recently research trends are changing to Transformer-based models. Although these Transformer-based models show good performance for time series data prediction problem, they show relatively insufficient performance for classification tasks. To address this problem, we propose a novel Transformer-based model to enhance the classification performance by adding CLS token to Transformer model and applying a pre-training method. The main contributions of this paper are summarized as follows : 1) an embedding method of input data, 2) a pre-training method for time series data. The embedding method of input data consists of two steps. First, the standard scaler is used to normalize time series data with different amplitudes. Then, we change the dimension of the data by using a time window method and use it as an input token for Transformer through GRU(Gated Recurrent Units). Second, we transform input data into an image using the GASF(Gramian Angular Summation Field). In addition, this transformed image is vectorized using a pre-trained model of ResNet. This vector is used as an input of a CLS token. Our pre-training method for time series data is basically based on the MLM (Masked Language Modeling) used in the natural language processing. Compared to the original MLM method, we use MSE(Mean Squared Error) for the evaluation of the objective function because time series data are composed of continuous variables unlike natural language processing. To show the efficacy of our method, we conduct extensive experiments with 12 different models using the UCR dataset. The experimental results show that our proposed model improves the average accuracy of 85 datasets from 1.4% to up to 21.1%.

    • Transformer-based Feature Extraction Approach for Hematopoietic Cancer Subtype Classification through Gene Expression Profile

      박광호 충북대학교 2024 국내박사

      RANK : 2943

      Recently, with the development of computing performance, machine learning, and deep learning, research is being actively conducted to solve various problems using the technology not only in the computer field but also in many other fields. As genetic information is becoming digitized in the field of biology, many studies are being conducted on ways to use computers to solve problems associated with various diseases using the genetic information. Cancer, which is mainly caused by genetic defects in cells, is one of the most active fields, and a lot of research has already been conducted. However, most studies are focused on classifying cancer and normal or cancers of different organs. The challenge is to clinically detect cancer in cells that have the ability to differentiate from a single cell into many different types of cells. Cells with this characteristic are called multipotent cells, and a typical cancer is breast cancer. In the case of breast cancer, the genetic markers of mammary stem cells are clear and have been utilized in various treatments. In contrast, hematopoietic stem cells, which are also multipotent cells, have the ability to differentiate into a wider variety of cells and differentiate in various locations in the body, making early diagnosis and prediction clinically difficult. Nevertheless, hematologic stem cell subtypes of hematologic cancers are less studied than other cancers, and unlike breast cancer, there are no accurate genetic markers of subtype differentiation. Therefore, this dissertation proposes a feature extraction technique using transformers to solve the subtype classification problem of hematopoietic cancer and detect genetic indicators. A transformer is a large structure that utilizes an attachment technique, which is currently being actively utilized in the field of natural language processing (NLP). In NLP, a transformer consists of an encoder and a decoder. The encoder is responsible for extracting contextual meaning, and the decoder is responsible for generating context from the extracted meaning. Using this concept, this dissertation proposes a transformer-based autoencoder (TFAE), a new feature extraction algorithm using a transformer encoder combined with autoencoder, with the goal of obtaining feature information meaningful for subtype classification of hematopoietic cancer. The details of research contents can be summarized as follows: First, this dissertation presents a transformer-based feature extraction algorithm, TFAE, for gene expression data. TFAE is designed to extract features by using a transformer-encoder to extract important feature of the data, and then extending it to a decoder to create the original. Second, in order to compare the feature extraction of the proposed feature extraction algorithm, different algorithms used for feature extraction are applied. For this purpose, PCA (Principal Component Analysis) and NMF (Non-Negative Factorization), which are widely used statistical-based feature extraction algorithms, and AE (Autoencoder) and VAE (Variational Autoencoder), which are deep learning-based feature extraction algorithms, were applied and compared. Third, multiple classifiers were applied to real tabular genomic data to classify hematopoietic cancer subtypes. Each set of features was applied to eight multiclass classifiers for performance evaluation. Finally, applied XAI (eXplainable Artificial Intelligence) to find genes that are important for hematopoietic cancer subtype classification. For this purpose, this research applied the SHAP (SHapley Additive exPlanations) algorithm, one of the XAI techniques, to show how much the extracted genes affect the subtyping and to detect which genes are important. To achieve these research objectives, the data of five types of blood cancers were collected from TCGA (The Cancer Genome Atlas), a representative open database of genetic information, and experimental data were generated by preprocessing, and then feature extraction algorithms including the proposed TFAE were used to extract genes of the same size. In order to determine how much a particular gene contributes to the subtype classification of hematopoietic cancers, this research applied the SHAP algorithm, one of the explanatory artificial intelligence (AI) techniques, to find the top 20 genes that best classify hematopoietic cancer subtypes. The overall experimental results show that the feature extraction techniques for each classifier yield reasonable performance for hematopoietic cancer subtype classification but the proposed TFAE algorithm can achieve better results than other feature extraction algorithms. In particular, when TFAE was combined with LGBM to classify hematopoietic cancer subtypes, the best performance was achieved with Accuracy 0.9857, Precision 0.9753, Recall 0.9635, Specificity 0.9963, F1 score 0.9691, G-mean 0.9797, and Balanced accuracy 0.9543. Although other algorithms have lower performance, they showed sufficiently significant performance in classification, confirming that this approach is effective. Consequently, the findings of this dissertation showed that our proposed feature extraction model, namely TFAE, could more accurately classify the hematopoietic cancer subtypes, and the SHAP method could identify the genes which are significant for each subtype classification. This dissertation can be regarded as one of the studies that showed the research potential of feature extraction techniques for classifier algorithms by applying transformer techniques to biological data, as apply real world hematopoietic cancer data to subtype classification. In the future, plans are to further develop this research and work on feature extraction for biological data using methods with similar representations such as Generative Adversarial Network (GAN) and diffusion.

    • Accelerating Transformer-Based Model Inference using Efficient Matrix Multiplications on GPUs

      이해룡 서울대학교 대학원 2023 국내박사

      RANK : 2943

      Transformer-based models have become the backbone of many state-of-the-art natural language processing (NLP) and computer vision tasks. As existing powerful models become large, enabling the models to learn and represent complex data relationships. Additionally, increasing the input sequence can be an effective way to improve performance for challenging real-world tasks. However, high inference cost hinders the use of powerful transformers because of large memory footprint, quadratic complexity with input sequence length in attention layers, and inefficient kernel operations. In this thesis, we propose Transformer optimization methods to reduce inference costs in various scenarios, depending on the model size, input sequence length, and batch size. First, we propose Multigrain, an optimization method for scenarios where the input length (Lin) is significantly greater than the hidden dimension (Dh). Existing sparse attention techniques can effectively reduce computation and memory footprints in long input sequences; however, they are inefficiently processed on GPUs and still account for the majority of the execution time. Multigrain takes into account the sparse patterns of sparse attention, processing the coarse-grained part with a coarse-grained kernel using high-performance tensor cores and the fine-grained part with a fine-grained kernel using CUDA cores, respectively. As a result, Multigrain achieves a 2.07x end-to-end speedup over DeepSpeed when running Longformer inference. Second, we propose a tiled singular value decomposition (TSVD) method to reduce inference costs in scenarios where Lin is similar to or smaller than Dh. TSVD is a technique that divides a matrix into tiles, performs singular value decomposition (SVD) on each tile, and compresses the matrix using low-rank approximation. By performing matrix multiplication, the fundamental operation of attention layers and feed-forward layers in Transformer models, using low-rank approximation-based TSVD-matmul, memory footprint and computation can be reduced, significantly lowering inference costs. Consequently, when compressing matrices by 2 to 8x, TSVD-based matrix multiplication is 1.02 to 2.26x faster than the uncompressed matrix multiplication. However, when applying TSVD to models, the execution time is reduced, but there is a trade-off in decreased accuracy. To address this issue, we propose TSVD-common, a parameter-efficient fine-tuning method based on TSVD. TSVD-common shares one of the submatrices decomposed by SVD in each tile across all tiles and fine-tunes only the common submatrix during training. As a result, TSVD-common improves accuracy by approximately 2% even when compressing the GPT-2 model by 2 or 4x in E2E NLG tasks, compared to full fine-tuning without compression. 최근 Transformer 기반의 모델들은 자연어 처리와 컴퓨터 비전 등 다양한 분야에서 높은 성능을 보여주고 있다. 기존 강력한 모델들은 커지면서 모델이 복잡한 데이터 관계를 학습하고 나타낼 수 있게 된다. 또한 입력 시퀀스 길이를 늘려 문맥학습을 향상시켜 복잡한 문제도 효과적으로 해결한다. 다만 이러한 모델들은 큰 메모리 사용량, 어텐션 레이어에서 입력 길이에 의한 2차복잡도 문제, 또한 커널 최적화가 되어 있지 않아 높은 추론 비용을 야기한다. 본 논문에서는 Transformer 기반 모델들의 크기, 입력 시퀀스 길이, 배치 크기에 따라 추론 비용을 줄이는 최적화 방법을 제안한다. 먼저, 입력 길이(Lin)가 은닉 차원(Dh)보다 큰 시나리오를 최적화하는 Multigrain 방법을 제안한다. 기존 희소 어텐션 기법은 긴 입력 시퀀스에서 연산량과 메모리 사용량을 효과적으로 줄일 수 있지만 GPU에서 비효율적으로 처리되며 여전히 대부분 수행시간을 차지한다. Multigrain은 희소 어텐션의 복합적인 희소 패턴을 파악하고 거친 희소 패턴은 고성능 텐서 코어를 사용한 커널로 처리하고 세밀한 패턴은 CUDA 코어를 사용한 커널로 각각 멀티 스트림으로 동시에 처리한다. 그 결과로 Longformer 모델을 DeepSpeed에서 추론을 실행한 기준 시스템에 비해 2.07배 더 빠른 것을 보여준다. 그리고 본 논문에서는 Lin이 Dh와 비슷하거나 작은 시나리오에서 추론 비용을 줄이는 tiled singular value decomposition(TSVD) 방법을 제안한다. TSVD는 행렬을 타일로 나누고 각 타일을 특이값 분해(SVD)하며 저랭크 근사를 이용하여 행렬을 압축하는 기법이다. Transformer 기반 모델에서 어텐션 레이어와 피드포워드 레이어의 기본 연산인 행렬 곱을 저랭크 근사를 이용한 TSVD기반의 행렬 곱으로 수행하면 메모리 사용량을 줄일 수 있고 연산량도 줄일 수 있으므로 추론 비용을 상당히 줄일 수 있다. 결과적으로 행렬을 2배~ 8배까지 압축 시, TSVD기반의 행렬 곱은 압축하지 않은 행렬 곱보다 1.02배--2.26배 빠른 것을 보인다. 다만 모델에 적용 시 수행시간이 줄어들지만 정확도가 하락하는 문제점이 존재한다. 이러한 문제점을 해결하기 위해 본 논문에서는 TSVD 기반의 매개변수 효율적 미세조정(parameter efficient fine-tuning) 방법인 TSVD-common을 제안한다. 각 타일에서 SVD로 분리된 두 서브행렬들 중 하나를 모든 타일에서 공유하는 형태로 하고 공동의 해당 서브행렬만 미세조정 시켜 학습시키는 방법이다. 결과적으로 제안한 TSVD-common은 GPT2 모델에서 2배 또는 4배 압축 시 E2E 태스크에서는 압축하지 않은 전체 매개변수를 미세조정하는 방법(full fine-tuning)보다 정확도가 2%정도 향상되었고 매개변수 효율적 미세조정 최신 방법인 LoRA와 근접한 정확도를 보여준다.

    • Transformer 기반 예측적 프로세스 자원 할당 알고리즘

      박영인 경기대학교 대학원 2023 국내석사

      RANK : 2943

      많은 기업들은 프로세스 기반 정보 시스템을 통해 업무 프로세스를 실행하고 관리한다. 프로세스 기반 정보 시스템은 실행된 프로세스 인스턴스의 이력을 프로세스 이벤트 로그 형태로 기록하고, 이러한 로그를 분석하여 효과적인 비즈니스 프로세스 운영을 지원한다. 프로세스 기반 정보 시스템을 사용하는 기업들은 프로세스 모니터링을 통해 커다란 경쟁 시장에서 우위를 선점하기 위해 업무 프로세스를 개선하고자 한다. 이에 따라 프로세스를 분석하고 개선하는 여러 연구들이 개발 되었으며, 특히 기업의 금전적인 손익과 직접적으로 연관이 되는 자원 할당에 대한 연구가 주목받고 있다. 이러한 연구는 실시간 프로세스 업무에 대하여 즉각적으로 자원을 할당하는 것에 대한 좋은 지표가 될 뿐만 아니라, 미래의 자원 할당 계획 수립에 도움이 된다. 프로세스 자원 할당 연구로는 규칙기반의 수학적 알고리즘과 자원의 적합성 및 관계파악, 자원 활용 모니터링이 가능한 프로세스 마이닝 등이 있다. 하지만 이들은 실시간으로 들어오는 데이터에 대해 예기치 못한 변수로 낮은 자원 활용률을 초래하거나, 업무 수행 흐름을 반영하지 못하여 비효율적인 자원 할당 계획을 수립할 수 있고, 자원의 상호 실행가능성을 고려하지 않아 프로세스 오작동을 초래할 수 있다. 본 논문에서는 미래 자원 할당 계획 수립을 위해 수행 흐름을 고려할 수 있는 예측적 프로세스 모니터링 기법으로 자원의 상호 실행가능성을 고려한 자원 할당 알고리즘을 제안한다. 예측적 프로세스 모니터링은 과거 프로세스 인스턴스의 실행이력을 분석하여 실행중인 프로세스 인스턴스의 미래 상태를 예측하는 기법으로, 업무 수행 흐름을 알 수 있는 다음 업무 및 런타임 예측이 가능하다. 최근 예측적 프로세스 모니터링에는 지능형 데이터 예측 기법인 딥러닝 기반 연구가 진행되었으며, 특히 순차적으로 기록되는 프로세스 이벤트 로그의 시계열적인 특성을 고려할 수 있는 시계열 딥러닝 모델인 LSTM을 사용한 연구가 가장 많이 이뤄졌다. 하지만 이는 프로세스 인스턴스의 길이가 길어질수록 성능이 낮아진다는 한계점이 있으며, 이에 따라 성능 향상을 위한 새로운 모델 기반의 연구가 필요했고, 이를 해결한 Transformer 기반의 연구가 이뤄졌다. Transformer는 기계 번역에서 SOTA 모델을 이룬 시계열 딥러닝 모델로, 출력을 예측하는 매 시점마다 전체 입력을 다시 참고함으로써 해당 출력과 연관성이 있는 입력에 대해 더 집중하는 Attention Mechanism을 사용한다. 프로세스 인스턴스는 업무의 흐름인 제어흐름을 가지며, 프로세스 인스턴스를 예측하는 데 있어 업무 간의 상관관계는 중요한 지표이다. Transformer는 Attention Mechanism을 통해 이러한 프로세스 업무 간의 상관관계를 고려한 예측이 가능하며, 이에 따라 Transformer를 사용한 예측적 프로세스 모니터링 연구인 ProcessTransformer가 등장하였다. ProcessTransformer는 다음 액티비티 예측에 있어 LSTM 기반 연구보다 훨씬 좋은 성능을 보였지만, 예측 모델의 입력으로 프로세스 액티비티 흐름만을 고려하였다. 프로세스 자원 할당 계획 수립을 위해서는 할당되었던 자원들에 대한 정보도 중요한 요소로 작용되며, 따라서 본 논문에서는 ProcessTransformer를 기반으로 자원의 정보도 고려하여 예측적 프로세스 모니터링 모델을 설계해 런타임 및 다음 업무를 예측하고, 자원의 상호 실행가능성을 고려하여 자원할당을 진행한다. 제안한 연구를 검증하고자 4TU.Centre for Research Data에서 제공하는 실제 프로세스 이벤트 로그인 Helpdesk, BPIC2012, BPIC2013, Review_Example_Large 데이터세트를 사용하여 실험한다. 예측적 프로세스 모니터링 성능 확인을 위해 기존 연구들과 비교한 결과, Helpdesk, BPIC2012, Review_Example_Large 데이터세트에서 가장 높은 성능을 보였다. 그리고 학습된 예측적 프로세스 모니터링 모델의 예측 결과를 통해 자원 할당을 하여 자원의 상호 실행가능성을 고려한 예측적 프로세스 인스턴스를 생성하였다. Many enterprises execute and manage business processes through Process-Aware Information System (PAIS). PAIS records the histories of executed process instances in the form of process event logs, and supports effective business process operation by analyzing these logs. Process-Aware enterprises seek to improve business processes to gain an advantage in a highly competitive market through process monitoring. Accordingly, several studies have been developed to analyze and improve the process, and in particular, research on resource allocation, which is directly related to financial profit-and-loss, is getting attention. These studies not only serves a good indicator of immediate resource allocation for current process tasks, but also helps to plan future resource allocation. Process resource allocation studies include rule-based mathematical algorithms and process mining, which can identify suitability and relationship of resources and monitor resource utilization. However, it can result in low resource utilization due to unexpected variables of data in real-time, and establish an inefficient resource allocation plan due to failing to reflect the workflow, and may cause malfunctions due to not consideting the interoperability of resources. For future resource allocation planning, this study proposes resource allocation algorithm considering interoperability of resources by performing Predictive Process Monitoring(PPM) that can consider the execution flow. PPM is a technique that predicts the future state of process instance, which is being executed, by analyzing the histories of past process instance and it is possible to predict th next task and runtime to figure out workflow. Recently, research based on deep learning, an intelligent data prediction technique, has been conducted for PPM, and in particular, LSTM, which can consider the time-series characteristics of sequentially recorded process event logs, was the most used. However, it has a limitation that the longer length of the process instance, the lower performance. Accordingly, a new model-based research was needed to improve performance, and Transformer-based research was conducted to solve this problem. Transformer is a time-series deep learning model that is state-of-the-art in machine translation, and uses Attention Mechanism that focuses more on inputs that are related to the output by referring back to the entire input at every point in predicting output. Process instance has a control flow, which is a workflow, and the correlation between tasks is an important indicator in predicting a process instance. Transformer can predict considering the correlation between process tasks through Attention Mechanism, and accordingly, ProcessTransformer, a PPM study using Transformer, appeared. ProcessTransformer performed significantly better than LSTM-based studies in predicting the next activity, but only considered the process activity flow as an input to predictive model. In order to establish a process resource allocation plan, information on allocated resources is also an important factor, therefore, this study is based on ProcessTransformer and design PPM model by considering resource factor to predict runtime and next task, and then proceed with resource allocation. To verify the proposed study, this paper experiments using the actual process event log datasets provided by 4TU.Centre for Research Data, Helpdesk, BPIC2012, BPIC2013, Review_Example_Large. As a result of comparison with previous studies to confirm PPM, proposed study showed high performance in Helpdesk, BPIC2012, and Review_Example_Large. Then, resource allocation was performed through the prediction result of the learned PPM model, and a predictive process instance was created considering interoperability of resources.

    • On the head redundancy in Swin transformer for image classification

      오준호 Graduate School, Yonsei University 2022 국내석사

      RANK : 2943

      Transformer 모델은 원래 자연어 처리를 위해 고안된 모델이지만, computer vision 등의 타 분야에서도 Transformer를 활용하려는 연구가 매우 활발하게 진행되고 있다. 자연어 처리 분야에서 이용되는 Transformer 기반 모델의 Multi-Head Self-Attention (MHSA) 모듈을 구성하는 각 헤드간의 중복성에 관한 연구는 존재하지만, computer vision 분야에서 이용되는 Transformer 기반 모델에 대해 이를 연구한 바는 없다. 본 논문에서는 Swin Transformer 모델의 MHSA 모듈에서 각 헤드를 제거했을 때 ImageNet 데이터셋의 이미지 분류 작업에서 모델 성능의 변화를 측정하였고, 이를 통해 몇몇 헤드는 다른 헤드와 중복성이 존재하여, 그 헤드를 제거하더라도 정확도의 하락이 적거나 오히려 상승함을 밝혔다. 또한 stage 3을 구성하는 헤드 중 절반을 제거하여 비교적 적은 정확도 감소를 대가로 모델을 경량화할 수 있음을 보였다. 마지막으로, 헤드의 중복성에 영향을 미칠 것으로 예상되는 3개의 잠재요소로 출력 행렬의 노름의 평균값, 입력에 따른 출력 행렬 간의 불변성, 각 헤드 간 attention map의 코사인 유사도를 제시하였고, 그 중 전자 2개의 요소와 모델 성능 간에 상관관계가 존재함을 발견하였다. Although the Transformer model was originally designed for natural language processing (NLP), many studies are being actively conducted in utilizing the Transformer in other fields such as computer vision. There is a study on the redundancy between attention heads in multi-head self-attention (MHSA) modules of Transformer-based models in NLP, but there is no study on Transformer-based models in computer vision. In this thesis, we measure the change of model performance on the ImageNet image classification task when each head in MHSA layers in Swin Transformer is removed. Several heads are redundant, so the model accuracy slightly decreases or rather increases. In addition, it is shown that the model can be compressed by removing half of the heads in stage 3, in exchange for insignificant accuracy loss. Finally, we offer three factors that are expected to affect the head redundancy; the mean of norms of output matrices for various inputs, the invariance of output matrices for various inputs, and the cosine similarity between attention maps of each head. It is proved that the former two factors are correlated with head redundancy.

    • TransSkip : Leveraging Transformer based Attention Mechanisms for Multi-Scale Feature Fusion

      Khan Rabeea Fatma 경기대학교 대학원 2024 국내석사

      RANK : 2943

      의료 영상 분할 분야의 효과적이고 정확한 해결책을 찾기 위해 다양한 연구가 진행되어 왔다. 전통적인 영상처리 방법에서 복잡한 합성곱 신경망(CNN)까지 다양한 방법이 제안되었다. 최근, Transformer의 셀프어텐션 메커니즘 및 Transformer와 CNN의 조합에서 비롯된 혼합 네트워크가 주된 방법으로 떠올랐다. 일반적으로 해당 모델은 계층적 인코더-디코더 아키텍처로 구성되어 있다. 이런 구조들은 간단한 잔차 연결을 사용하여 인코더에서 디코더로 중요한 공간 정보를 전송한다. 본 논문에서는 어떤 계층적 인코더 디코더 네트워크에서도 간단한 잔차 연결을 대체할 수 있는 TransSkip이라는 새로운 트랜스포머 기반 잔차 연결 아키텍처를 제안한다. 이는 교차 어텐션과 다중 해상도 상관 관계를 활용하여 전역 종속성 및 다양한 해상도에서의 정보를 효과적으로 포착한다. 간단한 잔차 연결을 제안된 향상된 스킵 연결로 대체함으로써 안정적인 네트워크를 구축하여 고도로 다양한 의료 데이터를 처리할 수 있다. 또한 TransSkip을 통해 최신 기술의 네트워크를 강화할 때 성능이 명확하게 향상되는 것을 방대한 실험 결과를 통해 확인할 수 있다. Significant research has been conducted to create efficient and accurate solutions to the very in demand problem of medical image segmentation. From handcrafted solutions to complex convolutional neural networks (CNNs), a variety of approaches have been looked into. One of these is the self-attention mechanism of the Transformer and the resulting hybrid networks of a combination of CNNs and the Transformer. Among these, the hierarchical encoder-decoder architecture has remained the state of the art. Generally, these models have simple skip connections transferring vital spatial information from the encoder to the decoder. In this thesis, a novel transformer-based skip connection architecture called TransSkip is proposed that can replace simple skip connections in any hierarchical encoder decoder network. This is carried out by utilizing cross attention and multi-scale correlations to efficiently capture global dependencies and information from different resolutions. Replacing simple skip connections with proposed enhanced skip connections can help construct robust networks that can deal with the highly varying medical data. This is also confirmed through extensive experimental results that show a clear boost in performance when state of the art networks are enhanced through TransSkip.

    • Reducing Computational Complexity of Transformer Models with U-Net-based Embedding Marvin John Ignacio

      IGNACIO MARVIN JOHN 세종대학교 대학원 2025 국내박사

      RANK : 2943

      Transformer-based models have become the foundation of modern artificial intel- ligence systems due to their expressive power and versatility in language, vision, and multimodal tasks. However, their widespread adoption remains limited by the high com- putational cost associated with training and deployment. This dissertation addresses the growing need for efficient model architectures by exploring a novel method to reduce the computational complexity of Transformer models through a U-Net-based embed- ding mechanism. We propose the U-Net Encapsulated Transformer (UET), an architectural framework that embeds Transformer modules within a hierarchical U-Net structure. By downsam- pling the token embeddings prior to self-attention, UET reduces the dimensionality of input representations, thereby minimizing the number of parameters and operations re- quired. This design enables efficient training from scratch, distinguishing it from most existing optimization methods that rely on post-hoc compression or pre-trained teacher models. UET is evaluated across three domains to demonstrate its cross-task generalizabil- ity. First, in multimodal sentiment analysis of internet memes, UET processes GPT- generated keyphrases and visual features to interpret nuanced emotional content with fewer resources. Second, in the sequential recommendation, UET constructs tempo- ral patterns in user behavior using a reinforcement learning framework, outperforming baselines on multiple datasets while preserving computational efficiency. Finally, UET is applied to language modeling, where it supports next-token prediction under reduced embedding dimensionality, offering a viable alternative to full-scale Transformers in low-resource environments. In addition to presenting this architectural solution, the dissertation integrates foun- dational concepts from the U-Net family of models. It reflects on the increasing im- portance of computational efficiency in user-facing applications. As multimodal and in- teractive AI systems become more prevalent, the demand for responsive and accessible models grows. UET responds to this demand by enabling practical, scalable deployment of Transformer-based systems without compromising depth or performance. This work contributes a general-purpose, resource-aware architecture for reducing the training and inference burden of Transformer models. Through theoretical motiva- tion and empirical validation, it offers a pathway toward more inclusive and sustainable AI development.

    • Refining Embeddings in Transformers : Image Classification, MRI Reconstruction, and Safe Image Generation = 트랜스포머 임베딩 정제 : 이미지 분류, MRI 재구성 및 안전한 이미지 생성

      안재신 경북대학교 대학원 2026 국내박사

      RANK : 2943

      Vision transformer research has evolved from foundational architectural designs to specialized domain applications and, more recently, to the critical issue of model safety. Following this trajectory, this thesis pursues three complementary lines of research to advance both the capability and responsible deployment of transformer models. First, to improve general-purpose vision backbones, I propose novel embedding architectures that utilize nonlinear transformations and shared-weight structures. Complementing these embedding-level improvements, a self-attention-based classifier head is introduced to replace the standard linear head, thereby improving feature aggregation at the final output stage. Second, adapting transformer architectures to the medical domain, I develop a domain-aware model for MRI reconstruction. This approach introduces a gated dual-domain transformer that processes spatial and frequency information concurrently to effectively correct off-resonance artifacts. Third, recognizing the safety risks in generative AI, I present a practical machine unlearning framework that controls embedding space. This method neutralizes unsafe concepts by realigning embeddings within the text encoder, ensuring robust defense against adversarial attacks while preserving benign generation quality. Collectively, these contributions demonstrate a comprehensive advancement of vision transformers, spanning from architectural refinement to domain specialization and safety alignment. 비전 트랜스포머 연구는 기초적인 아키텍처 설계에서 시작하여 특정 도메인으로의 응용, 그리고 최근에는 모델의 안전성 문제로 그 초점이 확장되어 왔다. 본 논문에서는 이러한 연구 흐름에 발맞추어 비전 트랜스포머의 성능과 활용성을 종합적으로 발전시키기 위한 세 가지 연구를 수행하였다. 첫째, 범용 비전 모델의 기초 역량을 강화하기 위해 임베딩 레이어와 분류기 구조를 개선하였다. 비선형 변환과 가중치 공유 구조를 도입하여 임베딩의 표현력을 높였으며, 기존의 선형 분류기를 대체하는 자기 주의 기반의 분류기 헤드를 도입하여 최종 단계에서의 특징 해석 성능을 개선하였다. 둘째, 의료 영상 분야의 특수한 요구사항을 반영하여 MRI 재구성을 위한 도메인 특화 트랜스포머를 개발하였다. 제안하는 게이트 이중 도메인 트랜스포머는 공간과 주파수 도메인 정보를 동시에 처리하여 MRI의 비공명 아티팩트를 효과적으로 보정한다. 셋째, 생성형 AI 모델의 안전한 활용을 위해 실용적인 머신 언러닝 프레임워크인 임베딩 공간 왜곡 기법을 제안하였다. 이 방법은 텍스트 인코더 내에서 유해 임베딩을 재정렬하여 생성 품질을 유지하면서도 유해 개념을 효과적으로 무력화한다. 결과적으로 본 논문은 아키텍처의 정교화부터 도메인 특화, 그리고 안전성 확보에 이르기까지 비전 트랜스포머 연구의 핵심적인 발전 방향을 포괄적으로 제시한다.

    • 트랜스포머 기반 동적 하이브리드 필터링과 위치 추정을 통한 UWB 측위 개선

      김현우 서강대학교 메타버스전문대학원 2025 국내석사

      RANK : 2943

      본 논문은 UWB(Ultra-Wideband) 기반 실내 위치 추정(Indoor Positioning System, IPS)에서 발생하는 NLOS(Non-Line of Sight) 문제와 다양한 노이즈 조건을 효과적으로 극복하기 위한 새로운 엔드투엔드 접근법을 제안한다. 기존 UWB 기반 IPS에서는 반사, 굴절 등으로 인한 신호 왜곡이 누적되어 오차가 증가하며, 특히 NLOS 환경에서 정확한 위치 추정이 어렵다는 한계가 존재한다. 이를 해결하기 위해 본 연구는 두 가지 핵심 요소를 결합한다: (1) 다양한 필터(KF, EKF, UKF, AKF, PF 등)를 하나의 풀(Pool)로 관리하고, 환경 변화에 따라 최적 필터를 실시간으로 선택하는 동적 하이브리드 필터링 프레임워크, (2) Transformer 기반 위치 추정 모델을 통한 복잡한 시계열 패턴 분석 및 안정적 위치 추정 수행이다. 전통적으로 실내 위치 추정에서는 특정 필터를 고정적으로 적용하거나 휴리스틱한 규칙에 따라 필터를 결정하는 방식이 주류를 이루었다. 그러나 이러한 접근법은 모든 상황에 대응하기 어렵고, NLOS 환경이나 다중 경로 반사 문제가 두드러지는 상황에서 성능 저하가 불가피하다. 반면 본 연구에서는 Transformer 아키텍처의 뛰어난 시계열 패턴 분석 능력을 활용하여 필터 선택을 데이터 기반 분류 문제로 재정의한다. 즉, Dynamic Hybrid Filtering Transformer를 통해 이전 시점의 추정 오차, TDOA(Time Difference of Arrival) 측정값 특성, IMU(가속도) 데이터 등을 분석하여, 현재 시점에 가장 적합한 필터를 선택함으로써 오차를 최소화한다. 이와 함께 시계열 데이터 처리에 특화된 Transformer 기반 위치 추정 모델을 적용하여, 단순히 순간값에 의존하지 않고 장기 의존성(Long-range dependency)을 학습한다. 이를 통해 NLOS 환경에서도 안정적으로 낮은 오차를 유지하며, Positional Encoding과 Time Series Adapter 모듈을 통해 시계열 특성을 강화함으로써 예측 정확도를 더욱 향상시킨다. 이런 방식은 각 시점의 노이즈 패턴, 다중 경로 반사 특성 등을 깊이 있게 파악하고, 필터를 동적으로 교체하여 최적의 결과를 도출할 수 있게 한다. 본 논문은 UTIL, Pfeiffer 공개 데이터셋을 활용하여 제안된 방법론의 성능을 검증하였으며, Baseline 대비 Basic Transformer 적용만으로도 유의미한 오차 감소가 확인되었다. 또한 개별 필터를 적용한 경우와 비교했을 때, 동적 하이브리드 필터링 전략을 도입하면 상황별 최적 필터를 선택함으로써 RMSE, MAE 등 위치 추정 오차 지표를 상당 폭 줄일 수 있음을 입증하였다. 나아가 Pfeiffer 데이터셋에서 일반화 성능을 검증함으로써 본 모델이 특정 환경에 과적합되지 않고 다양한 실내 환경 변화에도 유연하게 대응할 수 있는 가능성을 확인하였다. 본 연구의 기여점은 다음과 같다. 첫째, 필터 선택 문제를 Transformer 기반 데이터 주도적 접근으로 재정의하여, 정적인 필터 적용 방식의 한계를 극복했다. 둘째, Transformer 아키텍처를 통해 NLOS 환경에서의 비선형적 신호 패턴을 모델링함으로써, 뛰어난 강건성과 정확도를 확보했다. 셋째, 다양한 데이터셋을 통한 검증을 바탕으로 제안 프레임워크가 센티미터 단위 정확도가 요구되는 실내 위치 추적 문제에서 실용적인 솔루션이 될 수 있음을 제시했다. 향후 연구에서는 데이터셋 다양화, 모델 경량화, 온라인 학습 및 도메인 적응 기법 도입, 멀티센서 융합 등을 통해 제안된 프레임워크의 활용 범위를 넓히고 실시간성 및 설명 가능성을 강화할 수 있다. 이를 통해 스마트팩토리, 물류, 의료 시설, 로봇 공장 자동화 등 다양한 응용 분야에서 높은 정확도와 안정성을 갖춘 실내 위치 추정 서비스를 제공하는 핵심 기술로 발전시킬 수 있을 것으로 기대한다. This paper proposes a novel end-to-end approach to effectively address the Non-Line of Sight(NLOS) and noise-related challenges that arise in Ultra-Wideband(UWB) based Indoor Positioning Systems(IPS). While UWB theoretically offers centimeter-level precision, the presence of reflections, refractions, and complex indoor environments often leads to substantial inaccuracies, especially under NLOS conditions. To overcome these limitations, we combine two core strategies: (1) a dynamic hybrid filtering framework that manages a pool of diverse filters(KF, EKF, UKF, AKF, PF, etc.) and selects the most suitable filter in real-time according to environmental changes, and (2) a Transformer-based positioning model that exploits time-series pattern analysis for stable, accurate location estimation. Traditional IPS approaches often rely on a fixed filter or heuristic rules for filter selection, resulting in performance degradation under certain conditions. For instance, while Kalman-based filters may perform well in line-of-sight scenarios, they can falter in complex NLOS environments that demand more robust or nonlinear filtering strategies. In contrast, our research reframes filter selection as a data-driven classification problem solvable by a Dynamic Hybrid Filtering Transformer. This Transformer learns from previous estimation errors, TDOA (Time Difference of Arrival) characteristics, IMU data, and other environmental cues to predict, at each time step, which filter will minimize the current positioning error. Simultaneously, we employ a Transformer-based positioning model that takes advantage of multi-head attention, positional encoding, and a Time Series Adapter module to thoroughly capture long-range dependencies and complex temporal patterns in UWB signals. By integrating both these elements, we achieve robust performance even in environments plagued by NLOS effects and multipath reflections. Our approach reduces not only average errors but also improves system stability, ensuring that the model adapts rapidly to different patterns of noise and signal distortion. This paper verified the performance of the proposed methodology using the UTIL and Pfeiffer open datasets, and a significant error reduction was confirmed by applying the Basic Transformer alone compared to the Baseline. In addition, when compared to the case of applying individual filters, it was demonstrated that the introduction of a dynamic hybrid filtering strategy can significantly reduce RMSE, MAE, and other location estimation error indicators by selecting the optimal filter for each situation. Furthermore, by verifying the generalization performance on the Pfeiffer dataset, we confirmed the possibility that this model can flexibly respond to various indoor environment changes without overfitting to a specific environment. The key contributions of this study are as follows. First, we introduce a data-driven, Transformer-based method for optimal filter selection, surpassing the limitations of static or rule-based approaches. Second, we demonstrate that the Transformer’s time-series modeling capacity enables consistent accuracy and robustness, even under NLOS-induced nonlinearities and fluctuations. Third, we confirm these improvements through extensive validation on multiple datasets, reinforcing that the proposed framework is a versatile and practical solution for indoor localization tasks demanding high precision. In terms of future work, several avenues remain open. We can enhance generalization and adaptability by incorporating additional datasets and environments, explore model compression and optimization techniques for real-time or low-power devices, and consider online learning or domain adaptation strategies for continuous improvement. Furthermore, we may integrate other sensors (e.g., LiDAR, Wi-Fi, BLE) to enrich the data fusion capability of the model, thereby expanding its applicability to various industrial, logistics, healthcare, and navigation scenarios. Through these refinements, we aim to develop a robust, efficient, and widely applicable IPS solution that meets the stringent accuracy and reliability requirements of next-generation indoor positioning services.

    연관 검색어 추천

    이 검색어로 많이 본 자료

    활용도 높은 자료

    해외이동버튼