In this dissertation, robust cost functions for stereo matching are proposed to deal with occlusion, noise, and radiometric variation problems. Then, an efficient temporal data cost is presented for a real-time spatiotemporal stereo matching. A self-g...
In this dissertation, robust cost functions for stereo matching are proposed to deal with occlusion, noise, and radiometric variation problems. Then, an efficient temporal data cost is presented for a real-time spatiotemporal stereo matching. A self-guided cost aggregation is introduced to remove the dependency of a guidance image.
First, this research introduces a joint iterative anaglyph stereo matching and colorization framework for obtaining a set of disparity maps and colorized images. Conventional stereo matching algorithms fail when addressing anaglyph images that do not have similar intensities on their two respective view images (radiometric variation). To resolve this problem, two novel data costs using local color prior and reverse intensity distribution factor are proposed for obtaining accurate depth maps. To colorize an anaglyph image, each pixel in non-occluded regions is warped to another view using the obtained disparity value. A colorization algorithm using least square optimization is then employed with additional constraint to colorize the remaining occluded regions. Experimental results confirm that the proposed unified framework is robust and produces accurate depth maps and colorized stereo images. The generated depth map and color image for Teddy data obtains 7.24% bad pixels and 31.47 PSNR value which are the best among others.
Second, a light field depth estimation method that is more robust against occlusion and less sensitive to noise is introduced. Two novel data costs are proposed, which are measured using the angular patch and refocus image, respectively. The constrained angular entropy cost (CAE) reduces the effect of the dominant occluder and noise in the angular patch. The constrained adaptive defocus cost (CAD) provides a low cost in the occlusion region, while maintaining robustness against noise. Integrating the two data costs is shown to significantly improve the occlusion and noise invariant capability. Cost volume filtering and graph cut optimization are applied to improve the accuracy of the depth map. Experimental results confirm the robustness of the proposed method and demonstrate its ability to produce high-quality depth maps from a range of scenes. The proposed method outperforms other state-of-the-art light field depth estimation methods in both qualitative and quantitative evaluations. The average mean square error among the light field dataset is 0.0075 which is the smallest error among others.
Third, a novel Kinect-stereo camera fusion is presented to obtain a spatiotemporal consistent depth map video. Kinect and stereo camera suffers from inaccurate or missing depth information. To solve such problems, the proposed system builds a fusion camera that combines Kinect depth camera and stereo RGB camera. Both depth and disparity maps are efficiently integrated in a spatiotemporal MRF optimization function on a GPU. A novel temporal data cost is proposed to perform efficient depth map estimation while preserving the temporal coherency. Experimental results confirm that the proposed method is efficient, robust and accurate on challenging real-world data. The average temporal error of the proposed method in a dataset is 3.12%.
Finally, a deep self-guided cost aggregation method is presented used to obtain an accurate disparity map from a pair of stereo images. Conventional cost aggregation methods typically perform joint image filtering on each cost volume slice. Thus, a guidance image is necessary for the conventional methods to work effectively. However, a guidance image might be unreliable due to several distortions, such as noise, blur, radiometric variation.
To solve this problem, an advanced deep learning technique is used to perform self-guided cost aggregation. Because of the absence of ground truth cost volume, the solution for the dataset generation is offered. The proposed deep learning network consists of two sub-networks: dynamic weight network and descending filtering network. The feature reconstruction loss and the pixelwise mean square loss function are integrated to preserve the edge property. Experimental results show that the proposed method achieves better results even though it does not employ a guidance image. The average bad pixels percentage of Middlebury benchmark dataset is 17.06%.
In conclusion, this dissertation focuses on developing robust data costs for various problems in conventional stereo matching, such as occlusion, noise, radiometric variation, temporal consistency, etc. Each data cost is designed based on the observation of each input image type.
To handle the radiometric variation, it utilizes anaglyph image that represents a stereo image with large radiometric variation. Light field image is utilized to solve occlusion and noise problems because it contains lots of information of the corresponding pixels. Kinect depth is integrated with the stereo disparity map to handle difficult region, such as repetitive pattern. Then, temporal data cost is designed to achieve real-time performance in video stereo matching.
By utilizing the deep learning, the proposed cost aggregation method can utilize the input cost volume slice as the guidance. Therefore, the performance of cost aggregation is not affected by the input color image that may unreliable due to occlusion, noise or radiometric variation.