RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    계층적 FiLM 메커니즘 기반 소아 복부 종양 멀티모달 CT 영상 분할 모델 연구 : FiLM 주입 구조에서의 입력 모달리티 조합 및 파라미터 효율적 조건 벡터 결합 방식에 관한 절제 연구 = Study of a Multimodal CT Image Segmentation Model for Pediatric Abdominal Tumors Based on Hierarchical FiLM Mechanism

    한글로보기

    https://www.riss.kr/link?id=T17553702

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
      • URL 복사
    • 오류접수
    인용문이 복사되었습니다.

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Accurate image segmentation of pediatric abdominal masses remains a challenging task due to complex difficulties, including the rarity of the masses, anatomical variations across a wide age range, heterogeneity of multi-center imaging protocols, and ambiguous boundaries with surrounding organs. Recently, it has been reported that multimodal deep learning approaches combining non-image clinical information with image features can improve segmentation performance. However, systematic comparative studies on how to inject conditional information into the image feature space and how to fuse feature representations of heterogeneous data have rarely been conducted in the field of pediatric masses.
    This thesis proposes FusionSwinUNETRv2, a multimodal segmentation framework that adopts SwinUNETRv2 as the image encoder backbone while implementing a multi-resolution fusion structure, extending the predecessor study, FusionSwinUNETR, in two methodological directions. First, instead of the single-point structure that injected conditional information only into the first stage of the encoder via cross-attention, Feature-wise Linear Modulation (FiLM) blocks are applied to all four stages of the encoder. This ensures that patient-level non-image features modulate image features across all resolution levels of the hierarchical encoder. Second, to combine a 17-dimensional metadata vector and a 1024-dimensional clinical note embedding—obtained from OpenAI’s text-embedding-3-large model—into a single 256-dimensional conditional vector, three fusion strategies with varying adaptability and parameter scales are implemented as modules and directly compared: simple Summation (Sum), Adaptive Fusion with element-wise gating (Gated), and Concatenation followed by Projection (Concat).
    Experiments were conducted on a cohort of 750 pediatric patients with 12 types of masses from three tertiary general hospitals in South Korea, split into a 3-fold cross-validation dataset, resulting in a total of 60 independent test cases. The first experimental axis varied the modality combinations (Image, Image+Metadata, Image+Text, and Image+Metadata+Text) while fixing the fusion method to the element-wise Gated strategy. The second axis varied the fusion strategies (Sum, Gated, Concat) while fixing the modality combination to Image+Metadata+Text. Performance metrics included the Dice Similarity Coefficient (DSC), Precision, Recall, and the 95th percentile Hausdorff Distance (HD95).
    The model applying the simple Summation strategy, which has no learnable parameters for the three modalities, achieved the best overall performance with a DSC of 0.8952 ± 0.0475, Precision of 0.9088 ± 0.0739, Recall of 0.8911 ± 0.0792, and HD95 of 6.09 ± 4.09. It outperformed models using other fusion methods in three of the four metrics (DSC, Recall, and HD95). Furthermore, compared to the image-only baseline model (DSC 0.8825 ± 0.0674, HD95 6.70 ± 5.09), it demonstrated an average improvement of +1.27%p in DSC and a reduction of 0.61 in HD95.
    The result that simple Summation, lacking any learnable parameters, outperformed the learnable fusion methods (Gated and Concat) aligns with the performance increment patterns observed in the cross-attention-based model of the predecessor study. This provides important methodological insights regarding the balance between the data scale of this study and the number of parameters in the backbone model. This is interpreted to be because, in an architecture where modality-specific encoders already yield normalized conditional vectors through LayerNorm and GELU activations, simple Summation preserves the aligned scales of the modalities without causing training instabilities, such as the gate collapse phenomenon that can occur when passing through sigmoid gates. In other words, this implies that parameter efficiency and training stability, rather than the expressiveness of the fusion module itself, are the key factors determining the performance of the fusion strategies at this data scale.
    The contributions of this thesis are summarized in three points. First, to the best of our knowledge, this is the first study to conduct a two-way ablation experiment simultaneously controlling the conditional information injection architecture and fusion strategies for 12 types of pediatric abdominal masses. Second, it empirically confirms that the FiLM-based multi-stage injection architecture achieves average performance improvements compared to the SwinUNETRv2 baseline while maintaining parameter efficiency. Third, the finding that simple Summation outperforms more complex fusion methods provides the rationale for a practical design principle: "In data-scarce environments like pediatric mass imaging, keep the fusion module as simple as possible, and progressively add complexity in the form of residuals or pre-training during subsequent expansions." Future research directions include: exploring Residual Gated Fusion, which inherits the advantages of simple Summation while gradually introducing adaptability, or exploring pre-training strategies for the fusion module; validating generalizability by acquiring external public datasets or additional datasets from Gachon University Gil Medical Center; introducing systematic interpretability verification techniques such as Integrated Gradients; and refining training strategies, such as applying differential learning rates between the condition and image encoders, warmup schedules, various loss functions, and modality-specific dropouts.
    번역하기

    Accurate image segmentation of pediatric abdominal masses remains a challenging task due to complex difficulties, including the rarity of the masses, anatomical variations across a wide age range, heterogeneity of multi-center imaging protocols, and a...

    Accurate image segmentation of pediatric abdominal masses remains a challenging task due to complex difficulties, including the rarity of the masses, anatomical variations across a wide age range, heterogeneity of multi-center imaging protocols, and ambiguous boundaries with surrounding organs. Recently, it has been reported that multimodal deep learning approaches combining non-image clinical information with image features can improve segmentation performance. However, systematic comparative studies on how to inject conditional information into the image feature space and how to fuse feature representations of heterogeneous data have rarely been conducted in the field of pediatric masses.
    This thesis proposes FusionSwinUNETRv2, a multimodal segmentation framework that adopts SwinUNETRv2 as the image encoder backbone while implementing a multi-resolution fusion structure, extending the predecessor study, FusionSwinUNETR, in two methodological directions. First, instead of the single-point structure that injected conditional information only into the first stage of the encoder via cross-attention, Feature-wise Linear Modulation (FiLM) blocks are applied to all four stages of the encoder. This ensures that patient-level non-image features modulate image features across all resolution levels of the hierarchical encoder. Second, to combine a 17-dimensional metadata vector and a 1024-dimensional clinical note embedding—obtained from OpenAI’s text-embedding-3-large model—into a single 256-dimensional conditional vector, three fusion strategies with varying adaptability and parameter scales are implemented as modules and directly compared: simple Summation (Sum), Adaptive Fusion with element-wise gating (Gated), and Concatenation followed by Projection (Concat).
    Experiments were conducted on a cohort of 750 pediatric patients with 12 types of masses from three tertiary general hospitals in South Korea, split into a 3-fold cross-validation dataset, resulting in a total of 60 independent test cases. The first experimental axis varied the modality combinations (Image, Image+Metadata, Image+Text, and Image+Metadata+Text) while fixing the fusion method to the element-wise Gated strategy. The second axis varied the fusion strategies (Sum, Gated, Concat) while fixing the modality combination to Image+Metadata+Text. Performance metrics included the Dice Similarity Coefficient (DSC), Precision, Recall, and the 95th percentile Hausdorff Distance (HD95).
    The model applying the simple Summation strategy, which has no learnable parameters for the three modalities, achieved the best overall performance with a DSC of 0.8952 ± 0.0475, Precision of 0.9088 ± 0.0739, Recall of 0.8911 ± 0.0792, and HD95 of 6.09 ± 4.09. It outperformed models using other fusion methods in three of the four metrics (DSC, Recall, and HD95). Furthermore, compared to the image-only baseline model (DSC 0.8825 ± 0.0674, HD95 6.70 ± 5.09), it demonstrated an average improvement of +1.27%p in DSC and a reduction of 0.61 in HD95.
    The result that simple Summation, lacking any learnable parameters, outperformed the learnable fusion methods (Gated and Concat) aligns with the performance increment patterns observed in the cross-attention-based model of the predecessor study. This provides important methodological insights regarding the balance between the data scale of this study and the number of parameters in the backbone model. This is interpreted to be because, in an architecture where modality-specific encoders already yield normalized conditional vectors through LayerNorm and GELU activations, simple Summation preserves the aligned scales of the modalities without causing training instabilities, such as the gate collapse phenomenon that can occur when passing through sigmoid gates. In other words, this implies that parameter efficiency and training stability, rather than the expressiveness of the fusion module itself, are the key factors determining the performance of the fusion strategies at this data scale.
    The contributions of this thesis are summarized in three points. First, to the best of our knowledge, this is the first study to conduct a two-way ablation experiment simultaneously controlling the conditional information injection architecture and fusion strategies for 12 types of pediatric abdominal masses. Second, it empirically confirms that the FiLM-based multi-stage injection architecture achieves average performance improvements compared to the SwinUNETRv2 baseline while maintaining parameter efficiency. Third, the finding that simple Summation outperforms more complex fusion methods provides the rationale for a practical design principle: "In data-scarce environments like pediatric mass imaging, keep the fusion module as simple as possible, and progressively add complexity in the form of residuals or pre-training during subsequent expansions." Future research directions include: exploring Residual Gated Fusion, which inherits the advantages of simple Summation while gradually introducing adaptability, or exploring pre-training strategies for the fusion module; validating generalizability by acquiring external public datasets or additional datasets from Gachon University Gil Medical Center; introducing systematic interpretability verification techniques such as Integrated Gradients; and refining training strategies, such as applying differential learning rates between the condition and image encoders, warmup schedules, various loss functions, and modality-specific dropouts.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    소아 복부 종괴에 대한 정확한 영상 분할은 종괴의 희귀성, 광범위한 연령대에 따른 해부학적 변이, 다기관 영상 프로토콜의 이질성, 그리고 주변 장기와의 경계 모호성이라는 복합적 난제로 인해 여전히 도전적인 과제로 남아 있다. 최근 비영상 임상 정보와 영상 특징을 결합하는 멀티 모달 딥러닝 접근이 분할 성능을 향상시킬 수 있음이 보고되었으나, 조건 정보를 영상 특징 공간에 어떤 구조로 주입할 것인지, 그리고 이종의 데 이터에 대해 데이터별 특징를 어떻게 결합할 것인지에 대한 체계적 비교 연구는 소아 종괴 분야에서 거의 이루어지지 않았다. 본 학위논문은 SwinUNETRv2를 영상 인코더 백본으로 채택하면서도, 다 양한 해상도에 결합하는 구조를 구현한 멀티모달 분할 프레임워크인 FusionSwinUNETRv2를 제안하며, 본 연구의 선행 연구인 FusionSwinUNETR를 두 가지 방법론적 방향으로 확장한다. 첫째, 인코더 의 첫 번째 스테이지에만 교차 주의(cross-attention)를 통해 조건 정보를 주입하던 단일-지점 구조 Feature-wise Linear Modulation(FiLM) 블록을 인코더의 네 스테이지 모두에 적용함으로써, 환자 수준의 비영상 특징이 계층적 인코더의 모든 해상도 수준에서 영상 특징을 변조하도록 설계한 다. 둘째, 확보된 메타데이터와 텍스트 데이터를 기반으로 17차원 메타데 이터 벡터와 OpenAI사의 텍스트 임베딩 모델인 text-embedding-3-large 모델로부터 얻은 1024차원 임상 소견 임베딩을 단일 256차원 조건 벡터로 결합하기 위해 적응성과 파라미터 규모가 서로 다른 세 가지 결합 전략인 단순 합산(Sum), 원소 단위 게이트를 적용한 적응적 결합(Gated), 접합 후 투영(Concat)전략을 모듈 형태로 구현하여 직접 비교한다. 본 실험은 국내 3개 상급 종합병원의 12종의 종괴가 포함된 750명의 소 아 환자 코호트를 3겹 교차검증(3-fold cross-validation) 데이터셋으로 분 할하여 총 60건의 독립 테스트 사례에 대해 수행하였다. 첫 번째 실험은 결합 방식을 원소 단위 게이트를 적용한 적응적 결합으로 고정한 채 모달 리티 조합을 영상, 영상+메타데이터, 영상+텍스트데이터, 영상+메타데이터 +텍스트데이터로 변화시켰고, 두 번째 축은 모달리티 조합을 영상+메타데 이터+텍스트데이터로 고정한 채 결합 방식(Sum, Gated, Concat)을 변화시 켰다. 성능 지표로는 주사위 상관 계수(Dice Similarity Coefficient, DSC), 정밀도(Precision), 재현율(Recall), 95백분위 하우스도르프 거리 (95th percentaile Hausdorff distance, HD95)를 사용하였다. 세 모달리티에 학습 파라미터가 전혀 없는 단순 합산 방식을 적용한 모 델이 주사위 상관 계수 0.8952 ± 0.0475, 정밀도 0.9088 ± 0.0739, 재현 율 0.8911 ± 0.0792, 95백분위 하우스도르프 거리 6.09 ± 4.09로 가장 우 수한 종합 성능을 기록하였으며, 다른 결합 방식의 모델 대비 네 지표 중 세 지표(DSC, Recall, HD95)에서 우위를 보였다. 또한 영상 단일 모달리티 의 기준모델(DSC 0.8825 ± 0.0674, HD95 6.70 ± 5.09) 대비 평균 주사위 상관 계수가 +1.27%p, 95백분위 하우스도르프 거리가 −0.61 만큼 개선되 었다. 학습 파라미터가 전혀 없는 단순 합산 방식이 학습 가능한 결합 방식인 원소 단위 게이트를 적용한 적응적 결합 및 접합 후 투영 결합 방식을 상 회하였다는 결과는 선행 연구의 교차 주의 메커니즘 기반 모델에서의 모 달리티 누적에 따른 성능 증가 패턴과 유사하며, 본 연구의 데이터 규모 와 백본 모델 파라미터 수의 균형 관점에서 중요한 방법론적 통찰을 제공 한다. 이는 모달리티별 인코더가 이미 LayerNorm·GELU 활성화를 거쳐 정규화된 조건 벡터를 산출하는 구조에서는 단순 합산 방식이 두 모달리 티의 정렬된 스케일을 그대로 보존하며, 시그모이드 게이트를 거치는 과 정에서 발생하는 게이트 붕괴 현상 등과 같이 학습을 불안정하게 만들지 않기 때문으로 해석된다. 즉, 결합 모듈의 표현력 자체가 아니라 파라미 터 효율성과 학습 안정성이 본 데이터 규모에서 결합 전략의 성능을 결정 하는 핵심 요인임을 시사한다. 본 학위논문의 기여는 세 가지로 요약된다. 첫째, 소아 복부 종괴 12종 을 대상으로 조건 정보 주입 구조와 결합 전략을 동시에 통제한 이원적 절제 실험 연구를 본 연구가 알기로 처음으로 수행하였다. 둘째, FiLM 기 반 다단계 주입 아키텍처가 SwinUNETRv2 기준으로 대비되는 평균 성능 개선을 달성하면서도 파라미터 효율성을 유지함을 실증적으로 확인하였 다. 셋째, 단순 합산 방식이 다른 복잡한 결합 방식을 상회한 결과는 "소 아 종괴 영상과 같이 데이터 확보가 제한적인 환경에서는 결합 모듈을 가 능한 단순하게 유지하고, 이후 확장 시 잔차 또는 사전 학습 형태로 점진 적으로 추가하라"는 실용적 설계 원칙의 근거를 제공한다. 후속 연구로는 결합 모듈의 개선으로서 단순 합산의 이점을 계승하면서 적응성을 점진적 으로 도입하는 잔차 게이트 결합(Residual Gated Fusion)이나 결합 모듈의 사전 학습 전략을 탐색, 외부 공개 데이터셋 또는 가천대 길병원내 추가 데이터셋 확보를 통한 일반화 검증, 통합 기울기 등 체계적 해석 검증 기 법을 도입, 조건 인코더와 영상 인코더의 차등 학습률, 워밍업 스케줄, 다양한 손실 함수 적용, 모달리티 별 드롭아웃 등 학습 전략의 정교화가 제시된다.
    번역하기

    소아 복부 종괴에 대한 정확한 영상 분할은 종괴의 희귀성, 광범위한 연령대에 따른 해부학적 변이, 다기관 영상 프로토콜의 이질성, 그리고 주변 장기와의 경계 모호성이라는 복합적 난제...

    소아 복부 종괴에 대한 정확한 영상 분할은 종괴의 희귀성, 광범위한 연령대에 따른 해부학적 변이, 다기관 영상 프로토콜의 이질성, 그리고 주변 장기와의 경계 모호성이라는 복합적 난제로 인해 여전히 도전적인 과제로 남아 있다. 최근 비영상 임상 정보와 영상 특징을 결합하는 멀티 모달 딥러닝 접근이 분할 성능을 향상시킬 수 있음이 보고되었으나, 조건 정보를 영상 특징 공간에 어떤 구조로 주입할 것인지, 그리고 이종의 데 이터에 대해 데이터별 특징를 어떻게 결합할 것인지에 대한 체계적 비교 연구는 소아 종괴 분야에서 거의 이루어지지 않았다. 본 학위논문은 SwinUNETRv2를 영상 인코더 백본으로 채택하면서도, 다 양한 해상도에 결합하는 구조를 구현한 멀티모달 분할 프레임워크인 FusionSwinUNETRv2를 제안하며, 본 연구의 선행 연구인 FusionSwinUNETR를 두 가지 방법론적 방향으로 확장한다. 첫째, 인코더 의 첫 번째 스테이지에만 교차 주의(cross-attention)를 통해 조건 정보를 주입하던 단일-지점 구조 Feature-wise Linear Modulation(FiLM) 블록을 인코더의 네 스테이지 모두에 적용함으로써, 환자 수준의 비영상 특징이 계층적 인코더의 모든 해상도 수준에서 영상 특징을 변조하도록 설계한 다. 둘째, 확보된 메타데이터와 텍스트 데이터를 기반으로 17차원 메타데 이터 벡터와 OpenAI사의 텍스트 임베딩 모델인 text-embedding-3-large 모델로부터 얻은 1024차원 임상 소견 임베딩을 단일 256차원 조건 벡터로 결합하기 위해 적응성과 파라미터 규모가 서로 다른 세 가지 결합 전략인 단순 합산(Sum), 원소 단위 게이트를 적용한 적응적 결합(Gated), 접합 후 투영(Concat)전략을 모듈 형태로 구현하여 직접 비교한다. 본 실험은 국내 3개 상급 종합병원의 12종의 종괴가 포함된 750명의 소 아 환자 코호트를 3겹 교차검증(3-fold cross-validation) 데이터셋으로 분 할하여 총 60건의 독립 테스트 사례에 대해 수행하였다. 첫 번째 실험은 결합 방식을 원소 단위 게이트를 적용한 적응적 결합으로 고정한 채 모달 리티 조합을 영상, 영상+메타데이터, 영상+텍스트데이터, 영상+메타데이터 +텍스트데이터로 변화시켰고, 두 번째 축은 모달리티 조합을 영상+메타데 이터+텍스트데이터로 고정한 채 결합 방식(Sum, Gated, Concat)을 변화시 켰다. 성능 지표로는 주사위 상관 계수(Dice Similarity Coefficient, DSC), 정밀도(Precision), 재현율(Recall), 95백분위 하우스도르프 거리 (95th percentaile Hausdorff distance, HD95)를 사용하였다. 세 모달리티에 학습 파라미터가 전혀 없는 단순 합산 방식을 적용한 모 델이 주사위 상관 계수 0.8952 ± 0.0475, 정밀도 0.9088 ± 0.0739, 재현 율 0.8911 ± 0.0792, 95백분위 하우스도르프 거리 6.09 ± 4.09로 가장 우 수한 종합 성능을 기록하였으며, 다른 결합 방식의 모델 대비 네 지표 중 세 지표(DSC, Recall, HD95)에서 우위를 보였다. 또한 영상 단일 모달리티 의 기준모델(DSC 0.8825 ± 0.0674, HD95 6.70 ± 5.09) 대비 평균 주사위 상관 계수가 +1.27%p, 95백분위 하우스도르프 거리가 −0.61 만큼 개선되 었다. 학습 파라미터가 전혀 없는 단순 합산 방식이 학습 가능한 결합 방식인 원소 단위 게이트를 적용한 적응적 결합 및 접합 후 투영 결합 방식을 상 회하였다는 결과는 선행 연구의 교차 주의 메커니즘 기반 모델에서의 모 달리티 누적에 따른 성능 증가 패턴과 유사하며, 본 연구의 데이터 규모 와 백본 모델 파라미터 수의 균형 관점에서 중요한 방법론적 통찰을 제공 한다. 이는 모달리티별 인코더가 이미 LayerNorm·GELU 활성화를 거쳐 정규화된 조건 벡터를 산출하는 구조에서는 단순 합산 방식이 두 모달리 티의 정렬된 스케일을 그대로 보존하며, 시그모이드 게이트를 거치는 과 정에서 발생하는 게이트 붕괴 현상 등과 같이 학습을 불안정하게 만들지 않기 때문으로 해석된다. 즉, 결합 모듈의 표현력 자체가 아니라 파라미 터 효율성과 학습 안정성이 본 데이터 규모에서 결합 전략의 성능을 결정 하는 핵심 요인임을 시사한다. 본 학위논문의 기여는 세 가지로 요약된다. 첫째, 소아 복부 종괴 12종 을 대상으로 조건 정보 주입 구조와 결합 전략을 동시에 통제한 이원적 절제 실험 연구를 본 연구가 알기로 처음으로 수행하였다. 둘째, FiLM 기 반 다단계 주입 아키텍처가 SwinUNETRv2 기준으로 대비되는 평균 성능 개선을 달성하면서도 파라미터 효율성을 유지함을 실증적으로 확인하였 다. 셋째, 단순 합산 방식이 다른 복잡한 결합 방식을 상회한 결과는 "소 아 종괴 영상과 같이 데이터 확보가 제한적인 환경에서는 결합 모듈을 가 능한 단순하게 유지하고, 이후 확장 시 잔차 또는 사전 학습 형태로 점진 적으로 추가하라"는 실용적 설계 원칙의 근거를 제공한다. 후속 연구로는 결합 모듈의 개선으로서 단순 합산의 이점을 계승하면서 적응성을 점진적 으로 도입하는 잔차 게이트 결합(Residual Gated Fusion)이나 결합 모듈의 사전 학습 전략을 탐색, 외부 공개 데이터셋 또는 가천대 길병원내 추가 데이터셋 확보를 통한 일반화 검증, 통합 기울기 등 체계적 해석 검증 기 법을 도입, 조건 인코더와 영상 인코더의 차등 학습률, 워밍업 스케줄, 다양한 손실 함수 적용, 모달리티 별 드롭아웃 등 학습 전략의 정교화가 제시된다.

    더보기

    목차 (Table of Contents)

    • 제1장 서론 1
    • 1.1 연구 배경 및 목적 1
    • 1.2 연구 내용 및 방법 3
    • 제2장 소아 복부 종괴 세그멘테이션 연구 동향 6
    • 2.1 의료 영상 세그멘테이션 연구 동향 6
    • 제1장 서론 1
    • 1.1 연구 배경 및 목적 1
    • 1.2 연구 내용 및 방법 3
    • 제2장 소아 복부 종괴 세그멘테이션 연구 동향 6
    • 2.1 의료 영상 세그멘테이션 연구 동향 6
    • 2.1.1 소아 복부 종괴 영상 세그멘테이션의 임상적 난제 6
    • 2.1.2 단일 모달리티 기반 소아 종괴 분할 연구의 흐름 7
    • 2.1.3 멀티모달 의료 영상 AI 연구의 확산과 소아 분야의 공백 10
    • 2.2 핵심 기술 분석 12
    • 2.2.1 SwinUNETRv2 아키텍처 및 특징 12
    • 2.2.2 조건부 특징 결합 기법: Cross-Attention 과 FiLM 14
    • 2.2.3 이종 모달리티 결합 전략 비교 16
    • 제3장 멀티모달 FiLM 기반 세그멘테이션 모델 설계 19
    • 3.1 전체 모델 아키텍처 설계 19
    • 3.1.1 데이터 흐름 및 모듈 구성 20
    • 3.1.2 SwinUNETRv2 기반 영상 인코더 21
    • 3.2 조건부 벡터 생성 및 특징 주입 메커니즘 설계 23
    • 3.2.1 메타 인코더 및 텍스트 프로젝터 설계 23
    • 3.2.2 Fusion 모듈 설계 : Sum, Gated, Concat 25
    • 3.2.3 FiLM 기반 4-Stage 조건 주입 구조 27
    • 제4장 모델 구현 및 평가 31
    • 4.1 개발 환경 구축 31
    • 4.1.1 하드웨어 및 소프트웨어 환경 31
    • 4.1.2 데이터 구성 및 모델 학습 파이프라인 32
    • 4.2 핵심 기능 구현 35
    • 4.2.1 FusionSwinUNETRv2 모델 구현 35
    • 4.2.2 FiLM 블록 및 3종 Fusion 모듈 구현 36
    • 4.2.3 Grad-CAM 기반 3D 해석 가능성 모듈 37
    • 4.2.4 Ablation 자동화 파이프라인 37
    • 4.3 성능 및 해석 가능성 평가 38
    • 4.3.1 평가 메트릭 정의 38
    • 4.3.2 모달리티 조합 별 성능 비교 결과 39
    • 4.3.3 결합 방식 별 성능 비교 결과 41
    • 4.3.4 Grad-CAM 시각화 비교 분석 44
    • 제5장 분할 및 일반화 성능 증대 방안 47
    • 5.1 성능 및 결과 해석에 대한 고찰 47
    • 5.1.1 결과 요약 및 주요 발견 47
    • 5.1.2 외부 멀티모달 분할 연구와의 비교 고찰 48
    • 5.1.3 단순 합산이 우세했던 이유에 대한 분석 50
    • 5.2 제안한 모델의 한계 및 개선 방안 51
    • 5.2.1 Fusion 구조의 확장 가능성 51
    • 5.2.2 데이터 다양성 확보 및 외부 검증의 필요성 53
    • 5.2.3 해석 가능성 검증의 체계화 54
    • 5.2.4 학습 전략 및 최적화 관점의 개선 56
    • 제6장 결론 58
    • 참고문헌 61
    • ABSTRACT 68
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼