3D human modeling aims to reconstruct 3D human geometry from limited visual observations and to generate images or videos that are consistent with underlying 3D structure. Despite substantial progress, reliable human-centric 3D modeling remains challe...
3D human modeling aims to reconstruct 3D human geometry from limited visual observations and to generate images or videos that are consistent with underlying 3D structure. Despite substantial progress, reliable human-centric 3D modeling remains challenging due to the scarcity of high-quality 3D and 4D supervision, the articulated and non-rigid nature of the human body, and the increasing ambiguity encountered in stylized domains, multi-human scenes, and dynamic settings. While large-scale 2D image and video datasets are widely available, directly inferring or enforcing consistent 3D structure from such data remains fundamentally underconstrained without additional structural guidance.
This thesis studies geometry-aware diffusion models as a practical framework for human-centric 3D modeling under limited supervision. The central idea is to combine diffusion-based generative models, which learn strong appearance, structure, and motion representations from large-scale 2D data, with feasible geometric guidance such as pose, depth, surface normals, and visibility cues. These geometric signals serve to bridge 2D diffusion models and 3D structure, reducing ambiguity during generation and reconstruction.
The thesis is organized around three progressively more challenging scenarios. Part I addresses stylized and novel-domain 3D human and character generation, where target-domain geometric annotations are unavailable. It introduces diffusion-based pose-aware data generation and pose-preserved adaptation strategies that enable pretrained 3D generative models to adapt to new appearance domains while maintaining geometric consistency. Part II focuses on multi-human scenes, where occlusion, depth ordering ambiguity, and inter-person interaction complicate both generation and reconstruction. This part presents interaction-aware diffusion frameworks that incorporate occlusion-aware and group-level geometric representations to support coherent multi-human image synthesis and 3D reconstruction from limited visual input. Part III considers dynamic human scenes and studies temporally consistent 3D geometry estimation from monocular video. By reformulating dynamic geometry inference as a geometry-aware image-to-video diffusion process, this part leverages temporal priors learned by video diffusion models to stabilize geometry over time, reducing reliance on costly 4D supervision.
Overall, this thesis demonstrates that geometry-aware diffusion models provides a practical and extensible approach for human-centric 3D modeling under data scarcity. By selecting, controlling, and structuring geometric guidance according to the dominant source of uncertainty in each setting, the proposed framework enables coherent, controllable, and temporally stable 3D human generation and reconstruction across a wide range of challenging scenarios.