RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Learning Representations for Enhancing Multi-exposure High Dynamic Range Imaging = 다중 노출 고명암비 영상 복원을 위한 표현 학습 기법

    한글로보기

    https://www.riss.kr/link?id=T17314914

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Although image sensors have been advanced recently, most commercial cameras can still capture images with a limited range of illumination. Therefore, there is a growing need for high dynamic range (HDR) imaging techniques to acquire improved quality of the captured images. Specifically, multi-exposure HDR imaging shows better performance than single-image HDR imaging since multiple input low dynamic range (LDR) frames provide more information for compensating saturated regions. However, multi-exposure HDR imaging is a challenging task due to two major problems. One is misalignment between the input LDR frames and the other is missing information on LDR frames due to under-/over-exposed regions. Traditional multi-exposure HDR imaging methods introduced a frame merging approach with simple functions. These methods could produce high-quality HDR results when the LDR frames are captured with a strictly fixed camera and static environment, however, they did not consider the motions that occurred by camera or object movement. Recently, deep learning-based methods, specifically convolutional neural networks (CNNs), have shown notable improvement in various computer vision areas, including multi-exposure HDR imaging. However, these methods rely on end-to-end learning for training the network, neglecting the attributes of the LDR images with exposure difference and knowledge from the conventional exposure fusion process. To address these limitations, this dissertation introduces several novel approaches that effectively learns representative features of differently exposed LDR frames and HDR image representations.

    Firstly, a novel decomposition network is introduced that separately extracts content features and global features from LDR images. The content features help spatially align the input frames with exposure-invariant attribute, while global features, represented in the frequency domain using Fourier transforms, modulate the reconstruction process. An HDR reconstruction network further merges these features with multi-scale alignment and frequency domain modulation techniques. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across several benchmarks, significantly reducing ghosting and improving detail preservation compared to previous approaches.

    Additionally, a novel contrastive learning scheme for multi-exposure HDR imaging is proposed. This scheme introduces methods that synthesize LDR data pairs by defining the relationship between differently exposed LDR frames. The representation encoder is trained with those synthesized data pairs and contrastive objectives, and it is able to extract global and local representations from LDR frames. Furthermore, a novel HDR network that leverages learned representative features in the HDR reconstruction process, resulting in improved restoration performance.

    Finally, a novel overlapped codebook (OLC) is proposed for learning implicit HDR image representations. Specifically, the proposed approach employs a vector-quantization (VQ) mechanism and models the traditional exposure bracket process through its codebook structure. The OLC learns features of LDR frames with different exposure bias and HDR images concurrently, is able to represent HDR images with the combination of the LDR frames. A dual-decoder structured HDR network is also proposed for reconstructing high-quality HDR images, which leverages learned HDR representations to compensate severely saturated regions. Experimental results show that the proposed methods achieve state-of-the-art performance on diverse benchmarks and metrics.
    번역하기

    Although image sensors have been advanced recently, most commercial cameras can still capture images with a limited range of illumination. Therefore, there is a growing need for high dynamic range (HDR) imaging techniques to acquire improved quality o...

    Although image sensors have been advanced recently, most commercial cameras can still capture images with a limited range of illumination. Therefore, there is a growing need for high dynamic range (HDR) imaging techniques to acquire improved quality of the captured images. Specifically, multi-exposure HDR imaging shows better performance than single-image HDR imaging since multiple input low dynamic range (LDR) frames provide more information for compensating saturated regions. However, multi-exposure HDR imaging is a challenging task due to two major problems. One is misalignment between the input LDR frames and the other is missing information on LDR frames due to under-/over-exposed regions. Traditional multi-exposure HDR imaging methods introduced a frame merging approach with simple functions. These methods could produce high-quality HDR results when the LDR frames are captured with a strictly fixed camera and static environment, however, they did not consider the motions that occurred by camera or object movement. Recently, deep learning-based methods, specifically convolutional neural networks (CNNs), have shown notable improvement in various computer vision areas, including multi-exposure HDR imaging. However, these methods rely on end-to-end learning for training the network, neglecting the attributes of the LDR images with exposure difference and knowledge from the conventional exposure fusion process. To address these limitations, this dissertation introduces several novel approaches that effectively learns representative features of differently exposed LDR frames and HDR image representations.

    Firstly, a novel decomposition network is introduced that separately extracts content features and global features from LDR images. The content features help spatially align the input frames with exposure-invariant attribute, while global features, represented in the frequency domain using Fourier transforms, modulate the reconstruction process. An HDR reconstruction network further merges these features with multi-scale alignment and frequency domain modulation techniques. Extensive experiments demonstrate that the proposed method achieves state-of-the-art performance across several benchmarks, significantly reducing ghosting and improving detail preservation compared to previous approaches.

    Additionally, a novel contrastive learning scheme for multi-exposure HDR imaging is proposed. This scheme introduces methods that synthesize LDR data pairs by defining the relationship between differently exposed LDR frames. The representation encoder is trained with those synthesized data pairs and contrastive objectives, and it is able to extract global and local representations from LDR frames. Furthermore, a novel HDR network that leverages learned representative features in the HDR reconstruction process, resulting in improved restoration performance.

    Finally, a novel overlapped codebook (OLC) is proposed for learning implicit HDR image representations. Specifically, the proposed approach employs a vector-quantization (VQ) mechanism and models the traditional exposure bracket process through its codebook structure. The OLC learns features of LDR frames with different exposure bias and HDR images concurrently, is able to represent HDR images with the combination of the LDR frames. A dual-decoder structured HDR network is also proposed for reconstructing high-quality HDR images, which leverages learned HDR representations to compensate severely saturated regions. Experimental results show that the proposed methods achieve state-of-the-art performance on diverse benchmarks and metrics.

    더보기

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 이미지 센서들의 발전에도 불구하고, 대부분의 상용 카메라들은 여전히 제한된 조명 범위만을 촬영할 수 있다. 때문에 고품질의 촬영 이미지를 얻기 위해 하이 다이나믹 레인지 이미징 기술의 필요성이 증가하고 있다. 특히, 다중 노출 하이 다이나믹 레인지 이미징은 여러 장의 입력 로우 다이나믹 프레임의 정보들을 통해 포화 영역을 복원할 수 있기 때문에 단일 이미지 하이 다이나믹 레인지 이미징 기술보다 우수한 성능을 보여준다. 하지만 다중 노출 하이 다이나믹 레인지 이미징은 두개의 주된 문제 때문에 어려움을 겪는다. 하나는 입력 로우 다이나믹 레인지 프레임들이 정렬되지 않았다는 점과, 나머지는 로우 다이나믹 레인지 프레임들에 노출 부족 및 노출 과다로 인한 정보 손실이 있다는 점이다. 전통적인 다중 노출 하이 다이나믹 레인지 이미징 방법들은 단순한 함수를 통한 프레임 병합 방법을 제안하였다. 이러한 방법들은 고정된 카메라와 정적인 환경에서 촬영된 경우 고품질의 하이 다이나믹 레인지 이미지 결과를 제공할 수 있었으나, 카메라나 물체의 이동으로 발생하는 움직임을 고려하지 못하였다. 최근에는 딥 러닝 기반의 방법들, 특히 컨볼루션 신경망을 기반으로 하는 방법들이 다중 노출 하이 다이나믹 레인지 이미징을 포함한 다양한 컴퓨터 비전 분야에서 상당한 성능 향상을 보여주었다. 그러나 이러한 방법들은 엔드 투 엔드 학습을 통해 네트워크를 학습하며, 다양한 노출 정도의 로우 다이나믹 레인지 이미지들의 특성과 통상적인 노출 병합 과정을 반영하지 않는다. 이러한 문제를 해결하기 위해, 본 학위 논문에서는 다양한 노출 정도의 로우 다이나믹 레인지 이미지와 하이 다이나믹 레인지 이미지의 표현을 효과적으로 학습하는 방법을 제안한다.

    처음으로, 입력 로우 다이나믹 레인지 이미지를 콘텐트 피쳐와 전역 피쳐 두 피쳐로 분해하는 네트워크를 제안한다. 추출된 콘텐트 피쳐는 노출 차이에 불변하는 특징을 통해 입력 프레임을 정렬하는데 활용되며, 전역 피쳐는 푸리에 변환을 통해 주파수 영역에서 표현되어, 복원 과정에서 중간 피쳐를 조절하는 데 사용된다. 추출된 피쳐를 하이 다이나믹 레인지 복원 네트워크에 활용함으로써 다양한 특정 항목에서 기존 방법들보다 우수한 성능을 달성함을 보인다.

    다음으로, 다중 노출 하이 다이나믹 레인지 이미징을 위한 대조 학습 방법을 제안한다. 이 방법에서는 다양한 노출 정도의 로우 다이나믹 레인지 이미지들의 관계를 정의하고, 이에 따라 데이터 쌍을 생성하는 방법을 제안한다. 표현 인코더는 이렇게 생성된 데이터와 대조 학습 목표를 통해 학습되며, 로우 다이나믹 레인지 이미지로부터 전역 및 지역 표현을 추출할 수 있다. 더불어 추출된 표현을 복원 과정에 활용하여 우수한 성능을 보이는 새로운 하이 다이나믹 레인지 복원 네트워크를 제안한다.

    마지막으로, 하이 다이나믹 레인지 이미지의 암묵적 표현을 학습하는 새로운 오버랩 코드북 구조를 제안한다. 구체적으로 해당 방법은 벡터 양자화 메커니즘을 사용하여 전통적인 다중 노출 이미지 병합 과정을 코드북 구조를 통해 모델링한다. 제안한 오버랩 코드북은 다양한 노출 정도의 로우 다이나믹 레인지 이미지와 하이 다이나믹 레인지 이미지의 피쳐를 동시에 학습하며, 하이 다이나믹 레인지 이미지를 로우 다이나믹 레인지 이미지 표현의 조합으로 표현할 수 있다. 고품질의 하이 다이나믹 레인지 이미지를 복원 할 수 있는 듀얼 디코더 구조의 복원 네트워크 또한 제안하며, 이 네트워크는 사전 학습된 하이 다이나믹 레인지 이미지의 표현을 활용하여 포화된 영역을 보완한다. 실험 결과는 제안한 방법들이 다양한 벤치마크 데이터에 대해 다양한 측정항목에서 기존 방법들보다 우수한 성능을 달성함을 보인다.
    번역하기

    최근 이미지 센서들의 발전에도 불구하고, 대부분의 상용 카메라들은 여전히 제한된 조명 범위만을 촬영할 수 있다. 때문에 고품질의 촬영 이미지를 얻기 위해 하이 다이나믹 레인지 이미징...

    최근 이미지 센서들의 발전에도 불구하고, 대부분의 상용 카메라들은 여전히 제한된 조명 범위만을 촬영할 수 있다. 때문에 고품질의 촬영 이미지를 얻기 위해 하이 다이나믹 레인지 이미징 기술의 필요성이 증가하고 있다. 특히, 다중 노출 하이 다이나믹 레인지 이미징은 여러 장의 입력 로우 다이나믹 프레임의 정보들을 통해 포화 영역을 복원할 수 있기 때문에 단일 이미지 하이 다이나믹 레인지 이미징 기술보다 우수한 성능을 보여준다. 하지만 다중 노출 하이 다이나믹 레인지 이미징은 두개의 주된 문제 때문에 어려움을 겪는다. 하나는 입력 로우 다이나믹 레인지 프레임들이 정렬되지 않았다는 점과, 나머지는 로우 다이나믹 레인지 프레임들에 노출 부족 및 노출 과다로 인한 정보 손실이 있다는 점이다. 전통적인 다중 노출 하이 다이나믹 레인지 이미징 방법들은 단순한 함수를 통한 프레임 병합 방법을 제안하였다. 이러한 방법들은 고정된 카메라와 정적인 환경에서 촬영된 경우 고품질의 하이 다이나믹 레인지 이미지 결과를 제공할 수 있었으나, 카메라나 물체의 이동으로 발생하는 움직임을 고려하지 못하였다. 최근에는 딥 러닝 기반의 방법들, 특히 컨볼루션 신경망을 기반으로 하는 방법들이 다중 노출 하이 다이나믹 레인지 이미징을 포함한 다양한 컴퓨터 비전 분야에서 상당한 성능 향상을 보여주었다. 그러나 이러한 방법들은 엔드 투 엔드 학습을 통해 네트워크를 학습하며, 다양한 노출 정도의 로우 다이나믹 레인지 이미지들의 특성과 통상적인 노출 병합 과정을 반영하지 않는다. 이러한 문제를 해결하기 위해, 본 학위 논문에서는 다양한 노출 정도의 로우 다이나믹 레인지 이미지와 하이 다이나믹 레인지 이미지의 표현을 효과적으로 학습하는 방법을 제안한다.

    처음으로, 입력 로우 다이나믹 레인지 이미지를 콘텐트 피쳐와 전역 피쳐 두 피쳐로 분해하는 네트워크를 제안한다. 추출된 콘텐트 피쳐는 노출 차이에 불변하는 특징을 통해 입력 프레임을 정렬하는데 활용되며, 전역 피쳐는 푸리에 변환을 통해 주파수 영역에서 표현되어, 복원 과정에서 중간 피쳐를 조절하는 데 사용된다. 추출된 피쳐를 하이 다이나믹 레인지 복원 네트워크에 활용함으로써 다양한 특정 항목에서 기존 방법들보다 우수한 성능을 달성함을 보인다.

    다음으로, 다중 노출 하이 다이나믹 레인지 이미징을 위한 대조 학습 방법을 제안한다. 이 방법에서는 다양한 노출 정도의 로우 다이나믹 레인지 이미지들의 관계를 정의하고, 이에 따라 데이터 쌍을 생성하는 방법을 제안한다. 표현 인코더는 이렇게 생성된 데이터와 대조 학습 목표를 통해 학습되며, 로우 다이나믹 레인지 이미지로부터 전역 및 지역 표현을 추출할 수 있다. 더불어 추출된 표현을 복원 과정에 활용하여 우수한 성능을 보이는 새로운 하이 다이나믹 레인지 복원 네트워크를 제안한다.

    마지막으로, 하이 다이나믹 레인지 이미지의 암묵적 표현을 학습하는 새로운 오버랩 코드북 구조를 제안한다. 구체적으로 해당 방법은 벡터 양자화 메커니즘을 사용하여 전통적인 다중 노출 이미지 병합 과정을 코드북 구조를 통해 모델링한다. 제안한 오버랩 코드북은 다양한 노출 정도의 로우 다이나믹 레인지 이미지와 하이 다이나믹 레인지 이미지의 피쳐를 동시에 학습하며, 하이 다이나믹 레인지 이미지를 로우 다이나믹 레인지 이미지 표현의 조합으로 표현할 수 있다. 고품질의 하이 다이나믹 레인지 이미지를 복원 할 수 있는 듀얼 디코더 구조의 복원 네트워크 또한 제안하며, 이 네트워크는 사전 학습된 하이 다이나믹 레인지 이미지의 표현을 활용하여 포화된 영역을 보완한다. 실험 결과는 제안한 방법들이 다양한 벤치마크 데이터에 대해 다양한 측정항목에서 기존 방법들보다 우수한 성능을 달성함을 보인다.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents iii
    • List of Tables vi
    • List of Figures viii
    • 1 Introduction 1
    • Abstract i
    • Contents iii
    • List of Tables vi
    • List of Figures viii
    • 1 Introduction 1
    • 1.1 Contribution 4
    • 1.2 Contents 5
    • 2 Frequency Domain Multi-Exposure HDR Imaging Network with Representative Image Features 6
    • 2.1 Motivation and Overview 6
    • 2.2 Related Works 8
    • 2.2.1 Single-frame HDR imaging 9
    • 2.2.2 Multi-frame HDR imaging with Dynamic Scenes 9
    • 2.3 Proposed Methods 10
    • 2.3.1 Decomposition Network 10
    • 2.3.2 HDR Reconstruction Network 15
    • 2.4 Experimental Results 20
    • 2.4.1 Datasets 20
    • 2.4.2 Implementation Details 24
    • 2.4.3 Evaluation Metrics 25
    • 2.4.4 Comparison with the Previous Methods 25
    • 2.4.5 Analysis on Extracted Global Feature 27
    • 2.5 Ablation Study 29
    • 2.5.1 Computational Costs 29
    • 2.5.2 Impact of Proposed Modules 31
    • 2.5.3 Impact of Frequency-domain Loss 31
    • 2.6 Summary 32
    • 3 RFG-HDR: Representative Feature-Guided Transformer for Multi-exposure High Dynamic Range Imaging 34
    • 3.1 Motivation and Overview 34
    • 3.2 Related Work 36
    • 3.2.1 Multi-exposure HDR imaging 36
    • 3.2.2 Contrastive Learning 37
    • 3.3 Method 38
    • 3.3.1 Representative Feature Learning 38
    • 3.3.2 RFG-HDR 40
    • 3.4 Experiments 47
    • 3.4.1 Datasets 47
    • 3.4.2 Evaluation Metrics 47
    • 3.4.3 Implementation Details 47
    • 3.4.4 Comparison with Previous Methods 50
    • 3.5 Discussions 51
    • 3.5.1 Ablation Study 51
    • 3.5.2 Analysis of Global Representation 51
    • 3.5.3 Analysis of Local Representation 53
    • 3.5.4 Computational Costs 54
    • 3.6 Summary 54
    • 4 Enhancing Multi-Exposure High Dynamic Range Imaging with Overlapped Codebook for Improved Representation Learning 55
    • 4.1 Motivation and Overview 55
    • 4.2 Related Works 58
    • 4.2.1 Multi-Exposure HDR imaging 58
    • 4.2.2 Vector Quantization 59
    • 4.3 Method 59
    • 4.3.1 Overview 59
    • 4.3.2 Learning HDR representation with the OLC 60
    • 4.3.3 HDR imaging with learned representation 65
    • 4.4 Experiments 69
    • 4.4.1 Dataset 69
    • 4.4.2 Evaluation Metrics 70
    • 4.4.3 Implementation Details 70
    • 4.4.4 Quantitative Comparison 70
    • 4.4.5 Qualitative Comparison 73
    • 4.5 Discussion 73
    • 4.5.1 Analysis on the proposed OLC 73
    • 4.5.2 Impact of proposed modules 78
    • 4.6 Summary 79
    • 5 Conclusion 80
    • Abstract (In Korean) 92
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼