CUDA 기반 GPGPU 공유 메모리를 활용한 숄레스키 분해 알고리즘 구현 본 연구에서는 대칭 양의 정부호(SPD, Symmetric Positive Definite) 행렬의 효율적인 분해를 위해 CUDA 기반의 병렬 숄레스키 분해 알...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
CUDA 기반 GPGPU 공유 메모리를 활용한 숄레스키 분해 알고리즘 구현 본 연구에서는 대칭 양의 정부호(SPD, Symmetric Positive Definite) 행렬의 효율적인 분해를 위해 CUDA 기반의 병렬 숄레스키 분해 알...
CUDA 기반 GPGPU 공유 메모리를 활용한 숄레스키 분해 알고리즘 구현 본 연구에서는 대칭 양의 정부호(SPD, Symmetric Positive Definite) 행렬의 효율적인 분해를 위해 CUDA 기반의 병렬 숄레스키 분해 알고리즘 Hybrid Shared Memory Left-Looking/Right-Looking(HSMLL/HSMRL)을 제안한다. 위 알고리즘은 공유 메모리를 적극적으로 활용하여 데이터 재사용률을 높이고, 전역 메모리 접근 횟수를 최소화함으로써 GPU 내 통신 오버헤드를 감소시 키는 특징을 가지며, HSMLL 알고리즘은 패널 캐싱 구조, HSMRL 알고리즘 은 매크로타일 구조로 구성되어 행렬의 크기에 따라 적합한 알고리즘이 사 용되도록 구성했다. 실험 결과, RTX 4070 Ti Super GPU 상에서 제안한 알 고리즘은 기존 NVIDIA의 cusolverDnpotrf 커널에 비해 HSMLL 알고리즘은 행렬 크기가 512×512 이하인 소규모 연산에서 최대 약 1.20배의 가속화를, 1024×1024 이상의 대규모 행렬에서는 HSMRL 알고리즘이 최대 약 1.09배의 가속화를 보였다. 또한, Nsight Compute를 통한 프로파일링 분석에서 공유 메모리 활용률이 평균 92% 이상을 보이며 제안된 구조가 연산-메모리 균형 측면에서도 효율적임을 확인하였다.
다국어 초록 (Multilingual Abstract)
This study proposes a CUDA-based parallel Cholesky decomposition algorithm, Hybrid Shared Memory Left-Looking/Right-Looking (HSMLL/HSMRL), for efficient decomposition of symmetric positive definite (SPD) matrices. The algorithms exploit shared memory ...
This study proposes a CUDA-based parallel Cholesky decomposition
algorithm, Hybrid Shared Memory Left-Looking/Right-Looking
(HSMLL/HSMRL), for efficient decomposition of symmetric positive definite
(SPD) matrices. The algorithms exploit shared memory to enhance data
reuse and reduce global memory access, minimizing communication
overhead within the GPU. HSMLL employs a panel-caching structure
optimized for small matrices, while HSMRL adopts a macro-tile design
suitable for large matrices. Experiments on an NVIDIA RTX 4070 Ti Super
GPU show that HSMLL achieves up to 1.20× speedup for small matrices
(≤ 512×512), and HSMRL up to 1.09× for large matrices (≥ 1024×1024),
compared to NVIDIA cuSolverDnPotrf kernel. Nsight Compute profiling
reports over 92% shared memory utilization, confirming the proposed
approach achieves a balanced compute–memory performance.
목차 (Table of Contents)