RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Reducing Computational Complexity of Transformer Models with U-Net-based Embedding Marvin John Ignacio = U-Net 기반 임베딩을 활용한 트랜스포머 모델의 연산 복잡도 감소

    한글로보기

    https://www.riss.kr/link?id=T17287989

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Transformer-based models have become the foundation of modern artificial intel- ligence systems due to their expressive power and versatility in language, vision, and multimodal tasks. However, their widespread adoption remains limited by the high com- putational cost associated with training and deployment. This dissertation addresses the growing need for efficient model architectures by exploring a novel method to reduce the computational complexity of Transformer models through a U-Net-based embed- ding mechanism. We propose the U-Net Encapsulated Transformer (UET), an architectural framework that embeds Transformer modules within a hierarchical U-Net structure. By downsam- pling the token embeddings prior to self-attention, UET reduces the dimensionality of input representations, thereby minimizing the number of parameters and operations re- quired. This design enables efficient training from scratch, distinguishing it from most existing optimization methods that rely on post-hoc compression or pre-trained teacher models. UET is evaluated across three domains to demonstrate its cross-task generalizabil- ity. First, in multimodal sentiment analysis of internet memes, UET processes GPT- generated keyphrases and visual features to interpret nuanced emotional content with fewer resources. Second, in the sequential recommendation, UET constructs tempo- ral patterns in user behavior using a reinforcement learning framework, outperforming baselines on multiple datasets while preserving computational efficiency. Finally, UET is applied to language modeling, where it supports next-token prediction under reduced embedding dimensionality, offering a viable alternative to full-scale Transformers in low-resource environments. In addition to presenting this architectural solution, the dissertation integrates foun- dational concepts from the U-Net family of models. It reflects on the increasing im- portance of computational efficiency in user-facing applications. As multimodal and in- teractive AI systems become more prevalent, the demand for responsive and accessible models grows. UET responds to this demand by enabling practical, scalable deployment of Transformer-based systems without compromising depth or performance. This work contributes a general-purpose, resource-aware architecture for reducing the training and inference burden of Transformer models. Through theoretical motiva- tion and empirical validation, it offers a pathway toward more inclusive and sustainable AI development.
    번역하기

    Transformer-based models have become the foundation of modern artificial intel- ligence systems due to their expressive power and versatility in language, vision, and multimodal tasks. However, their widespread adoption remains limited by the high com...

    Transformer-based models have become the foundation of modern artificial intel- ligence systems due to their expressive power and versatility in language, vision, and multimodal tasks. However, their widespread adoption remains limited by the high com- putational cost associated with training and deployment. This dissertation addresses the growing need for efficient model architectures by exploring a novel method to reduce the computational complexity of Transformer models through a U-Net-based embed- ding mechanism. We propose the U-Net Encapsulated Transformer (UET), an architectural framework that embeds Transformer modules within a hierarchical U-Net structure. By downsam- pling the token embeddings prior to self-attention, UET reduces the dimensionality of input representations, thereby minimizing the number of parameters and operations re- quired. This design enables efficient training from scratch, distinguishing it from most existing optimization methods that rely on post-hoc compression or pre-trained teacher models. UET is evaluated across three domains to demonstrate its cross-task generalizabil- ity. First, in multimodal sentiment analysis of internet memes, UET processes GPT- generated keyphrases and visual features to interpret nuanced emotional content with fewer resources. Second, in the sequential recommendation, UET constructs tempo- ral patterns in user behavior using a reinforcement learning framework, outperforming baselines on multiple datasets while preserving computational efficiency. Finally, UET is applied to language modeling, where it supports next-token prediction under reduced embedding dimensionality, offering a viable alternative to full-scale Transformers in low-resource environments. In addition to presenting this architectural solution, the dissertation integrates foun- dational concepts from the U-Net family of models. It reflects on the increasing im- portance of computational efficiency in user-facing applications. As multimodal and in- teractive AI systems become more prevalent, the demand for responsive and accessible models grows. UET responds to this demand by enabling practical, scalable deployment of Transformer-based systems without compromising depth or performance. This work contributes a general-purpose, resource-aware architecture for reducing the training and inference burden of Transformer models. Through theoretical motiva- tion and empirical validation, it offers a pathway toward more inclusive and sustainable AI development.

    더보기

    목차 (Table of Contents)

    • I Introduction 1
    • 1 Background of the Study 1
    • 2 Optimization in UI/UX-Driven Generative AI 1
    • 3 Problem Statement 2
    • 4 Existing Approaches to Model Optimization 3
    • I Introduction 1
    • 1 Background of the Study 1
    • 2 Optimization in UI/UX-Driven Generative AI 1
    • 3 Problem Statement 2
    • 4 Existing Approaches to Model Optimization 3
    • 5 U-Net Encapsulated Transformer (UET) 3
    • 6 Research Contributions 4
    • 7 Dissertation Structure 5
    • II Related Work 7
    • 1 Existing AI Systems 7
    • 2 Transformer Architecture 8
    • 2.1 Self-Attention Mechanism 9
    • 2.2 Feed-Forward Network Variants 10
    • 2.3 Architectural Extensions 10
    • 3 Optimization Techniques 11
    • 4 U-Net and Its Architectural Enhancements 13
    • 4.1 Architectural Structure 14
    • 4.2 Architectural Enhancements and Advantages 16
    • 4.3 Model Integration with U-Net 22
    • III UET for Multimodal Sentiment Analysis 25
    • 1 Introduction 25
    • 2 Methodology 26
    • 2.1 Overview 26
    • 2.2 Image Captioning 28
    • 2.3 Context Generation 29
    • 2.4 Keyphrase Embeddings 29
    • 2.5 UET Model 30
    • 2.6 Classification Model 31
    • 3 Experiments and Results 31
    • 3.1 Dataset and Evaluation Metrics 31
    • 3.2 Implementation Details 32
    • 3.3 Ablation Study 33
    • 3.3.1 Effects of keyphrase embeddings 36
    • 3.4 Comparison with the challenge participants 37
    • IV UET for Sequential Recommendation 41
    • 1 Introduction 41
    • 2 Methodology 44
    • 2.1 Overall Architecture 44
    • 2.2 U-Net Encapsulated Transformer 45
    • 2.3 Model Augmentation 47
    • 2.4 Prediction Layer 48
    • 2.5 Objective Function 48
    • 2.5.1 Supervised Learning 49
    • 2.5.2 Reinforcement Learning 49
    • 2.5.3 Contrastive Learning 50
    • 2.5.4 Final Objective Function 51
    • 3 Experiments and Results 52
    • 3.1 Datasets 52
    • 3.2 Evaluation 52
    • 3.3 Implementation Details 53
    • 3.4 Ablation Study 54
    • 3.4.1 Effects of Encapsulating the Transformer 54
    • 3.4.2 Comparison with Existing Architectures 56
    • 3.4.3 Effect of Embedding Size 57
    • 3.4.4 Effect of Contrastive Learning 57
    • 3.5 Comparison with Baseline Models 58
    • V UET for Language Modeling 62
    • 1 Introduction 62
    • 2 Methodology 63
    • 2.1 Tokenization and Embedding Layer 63
    • 2.2 U-Net Encoder and Decoder Layers 64
    • 2.3 Transformer 65
    • 2.4 Objective Function 66
    • 2.5 Training Iterations and Batch Size Considerations 66
    • 3 Experiments and Results 66
    • 3.1 Ablation Study 67
    • 3.2 Language Modeling 72
    • VI Discussion 79
    • 1 Meme Sentiment Analysis 79
    • 2 Sequential Recommender 80
    • 3 Language Modeling 80
    • 4 Implications for UI/UX and Interactive Systems 81
    • VII Conclusion 83
    • BIBLIOGRAPHY 85
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼