Accurate image segmentation of pediatric abdominal masses remains a challenging task due to complex difficulties, including the rarity of the masses, anatomical variations across a wide age range, heterogeneity of multi-center imaging protocols, and a...
Accurate image segmentation of pediatric abdominal masses remains a challenging task due to complex difficulties, including the rarity of the masses, anatomical variations across a wide age range, heterogeneity of multi-center imaging protocols, and ambiguous boundaries with surrounding organs. Recently, it has been reported that multimodal deep learning approaches combining non-image clinical information with image features can improve segmentation performance. However, systematic comparative studies on how to inject conditional information into the image feature space and how to fuse feature representations of heterogeneous data have rarely been conducted in the field of pediatric masses.
This thesis proposes FusionSwinUNETRv2, a multimodal segmentation framework that adopts SwinUNETRv2 as the image encoder backbone while implementing a multi-resolution fusion structure, extending the predecessor study, FusionSwinUNETR, in two methodological directions. First, instead of the single-point structure that injected conditional information only into the first stage of the encoder via cross-attention, Feature-wise Linear Modulation (FiLM) blocks are applied to all four stages of the encoder. This ensures that patient-level non-image features modulate image features across all resolution levels of the hierarchical encoder. Second, to combine a 17-dimensional metadata vector and a 1024-dimensional clinical note embedding—obtained from OpenAI’s text-embedding-3-large model—into a single 256-dimensional conditional vector, three fusion strategies with varying adaptability and parameter scales are implemented as modules and directly compared: simple Summation (Sum), Adaptive Fusion with element-wise gating (Gated), and Concatenation followed by Projection (Concat).
Experiments were conducted on a cohort of 750 pediatric patients with 12 types of masses from three tertiary general hospitals in South Korea, split into a 3-fold cross-validation dataset, resulting in a total of 60 independent test cases. The first experimental axis varied the modality combinations (Image, Image+Metadata, Image+Text, and Image+Metadata+Text) while fixing the fusion method to the element-wise Gated strategy. The second axis varied the fusion strategies (Sum, Gated, Concat) while fixing the modality combination to Image+Metadata+Text. Performance metrics included the Dice Similarity Coefficient (DSC), Precision, Recall, and the 95th percentile Hausdorff Distance (HD95).
The model applying the simple Summation strategy, which has no learnable parameters for the three modalities, achieved the best overall performance with a DSC of 0.8952 ± 0.0475, Precision of 0.9088 ± 0.0739, Recall of 0.8911 ± 0.0792, and HD95 of 6.09 ± 4.09. It outperformed models using other fusion methods in three of the four metrics (DSC, Recall, and HD95). Furthermore, compared to the image-only baseline model (DSC 0.8825 ± 0.0674, HD95 6.70 ± 5.09), it demonstrated an average improvement of +1.27%p in DSC and a reduction of 0.61 in HD95.
The result that simple Summation, lacking any learnable parameters, outperformed the learnable fusion methods (Gated and Concat) aligns with the performance increment patterns observed in the cross-attention-based model of the predecessor study. This provides important methodological insights regarding the balance between the data scale of this study and the number of parameters in the backbone model. This is interpreted to be because, in an architecture where modality-specific encoders already yield normalized conditional vectors through LayerNorm and GELU activations, simple Summation preserves the aligned scales of the modalities without causing training instabilities, such as the gate collapse phenomenon that can occur when passing through sigmoid gates. In other words, this implies that parameter efficiency and training stability, rather than the expressiveness of the fusion module itself, are the key factors determining the performance of the fusion strategies at this data scale.
The contributions of this thesis are summarized in three points. First, to the best of our knowledge, this is the first study to conduct a two-way ablation experiment simultaneously controlling the conditional information injection architecture and fusion strategies for 12 types of pediatric abdominal masses. Second, it empirically confirms that the FiLM-based multi-stage injection architecture achieves average performance improvements compared to the SwinUNETRv2 baseline while maintaining parameter efficiency. Third, the finding that simple Summation outperforms more complex fusion methods provides the rationale for a practical design principle: "In data-scarce environments like pediatric mass imaging, keep the fusion module as simple as possible, and progressively add complexity in the form of residuals or pre-training during subsequent expansions." Future research directions include: exploring Residual Gated Fusion, which inherits the advantages of simple Summation while gradually introducing adaptability, or exploring pre-training strategies for the fusion module; validating generalizability by acquiring external public datasets or additional datasets from Gachon University Gil Medical Center; introducing systematic interpretability verification techniques such as Integrated Gradients; and refining training strategies, such as applying differential learning rates between the condition and image encoders, warmup schedules, various loss functions, and modality-specific dropouts.