Although image sensors have been advanced recently, most commercial cameras can still capture images with a limited range of illumination. Therefore, there is a growing need for high dynamic range (HDR) imaging techniques to acquire improved quality o...
Although image sensors have been advanced recently, most commercial cameras can still capture images with a limited range of illumination. Therefore, there is a growing need for high dynamic range (HDR) imaging techniques to acquire improved quality of the captured images. Specifically, multi-exposure HDR imaging shows better performance than single-image HDR imaging since multiple input low dynamic range (LDR) frames provide more information for compensating saturated regions. However, multi-exposure HDR imaging is a challenging task due to two major problems. One is misalignment between the input LDR frames and the other is missing information on LDR frames due to under-/over-exposed regions. Traditional multi-exposure HDR imaging methods introduced a frame merging approach with simple functions. These methods could produce high-quality HDR results when the LDR frames are captured with a strictly fixed camera and static environment, however, they did not consider the motions that occurred by camera or object movement. Recently, deep learning-based methods, specifically convolutional neural networks (CNNs), have shown notable improvement in various computer vision areas, including multi-exposure HDR imaging. However, these methods rely on end-to-end learning for training the network, neglecting the attributes of the LDR images with exposure difference and knowledge from the conventional exposure fusion process. To address these limitations, this dissertation introduces several novel approaches that effectively learns representative features of differently exposed LDR frames and HDR image representations.
Firstly, a novel decomposition network is introduced that separately extracts content features and global features from LDR images. The content features help spatially align the input frames with exposure-invariant attribute, while global features, represented in the frequency domain using Fourier transforms, modulate the reconstruction process. An HDR reconstruction network further merges these features with multi-scale alignment and frequency domain modulation techniques. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across several benchmarks, significantly reducing ghosting and improving detail preservation compared to previous approaches.
Additionally, a novel contrastive learning scheme for multi-exposure HDR imaging is proposed. This scheme introduces methods that synthesize LDR data pairs by defining the relationship between differently exposed LDR frames. The representation encoder is trained with those synthesized data pairs and contrastive objectives, and it is able to extract global and local representations from LDR frames. Furthermore, a novel HDR network that leverages learned representative features in the HDR reconstruction process, resulting in improved restoration performance.
Finally, a novel overlapped codebook (OLC) is proposed for learning implicit HDR image representations. Specifically, the proposed approach employs a vector-quantization (VQ) mechanism and models the traditional exposure bracket process through its codebook structure. The OLC learns features of LDR frames with different exposure bias and HDR images concurrently, is able to represent HDR images with the combination of the LDR frames. A dual-decoder structured HDR network is also proposed for reconstructing high-quality HDR images, which leverages learned HDR representations to compensate severely saturated regions. Experimental results show that the proposed methods achieve state-of-the-art performance on diverse benchmarks and metrics.