스케치 기반 애니메이션 얼굴 색채화는 구조 보존, 스타일 일관성, 사실성을 동시에 달성해야 하는 어려운 조건부 생성 문제이다. 기존 DDPM 및 EDM 기반 확산 모델은 cosine 등 비지각적 시그마 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
스케치 기반 애니메이션 얼굴 색채화는 구조 보존, 스타일 일관성, 사실성을 동시에 달성해야 하는 어려운 조건부 생성 문제이다. 기존 DDPM 및 EDM 기반 확산 모델은 cosine 등 비지각적 시그마 ...
스케치 기반 애니메이션 얼굴 색채화는 구조 보존, 스타일 일관성, 사실성을 동시에 달성해야 하는 어려운 조건부 생성 문제이다. 기존 DDPM 및 EDM 기반 확산 모델은 cosine 등 비지각적 시그마 스케줄로 인해 timestep별 복원 난이도가 불균형해지며, 특히 스케치, 참조 불일치 환경에서 구조 붕괴와 색 번짐이 빈번하게 발생한다. 본 연구는 AnimeDiffusion을 baseline으로 하여 모델 구조 변경 없이 DDPM 기반 프레임워크를 EDM 기반 σ 표현으로 전환하고, SSIM으로 계측한 지각 손상 난이도에 따라 노이즈 스케줄을 재정렬하는 시그마 스케일링 기법 SSIMBaD (Sigma Scaling with SSIM-Guided Balanced Diffusion)를 제안한다. 제안 기법은 학습/추론 단계 모두에서 σ를 균일하게 샘플링/스케줄링하도록 설계하여, 모든 난이도의 노이즈에 대해 학습된 복원 능력이 추론 시 일관되게 활용되도록 한다. 또한 역확산의 결과를 Finetuning에 활용한 trajectory refinement를 결합하여 구조적 안정성을 유지하면서 세부 색감과 스타일 정합성을 개선하였다. 실험결과 AnimeDiffusion, SCFT, ControlNet, train-free attention 기반 기법들 대비 구조, 스타일, 사실성 전반에서 더 균형 잡힌 성능을 보였으며, 특히 cross-reference 환경에서 FID와 MS-SSIM([34])을 유의미하게 향상시켰다. 본 연구는 SSIM 기반 σ 재정렬과 역경로 미세 조정이 구조 보존이 중요한 조건부 생성 문제 전반에 적용 가능한 일반적 설계 원리를 보여준다.
다국어 초록 (Multilingual Abstract)
Sketch-based anime face colorization is a challenging conditional generation task that must simultaneously preserve structural fidelity, maintain stylistic consistency, and achieve visual realism. Existing diffusion models based on DDPM and EDM typica...
Sketch-based anime face colorization is a challenging conditional
generation task that must simultaneously preserve structural
fidelity, maintain stylistic consistency, and achieve visual realism.
Existing diffusion models based on DDPM and EDM typically rely
on non-perceptual noise schedules, such as cosine schedules,
which lead to imbalanced reconstruction difficulty across
timesteps. This imbalance is particularly problematic in
sketch reference mismatch scenarios, where it often results in –
structural collapse and color bleeding artifacts. In this work, we
take AnimeDiffusion as a baseline and propose SSIMBaD (Sigma
Scaling with SSIM-Guided Balanced Diffusion), a sigma scaling
method that reorganizes the noise schedule according to
perceptual corruption difficulty measured by SSIM, without
modifying the underlying model architecture. Specifically, we
replace the conventional DDPM-based formulation with an
EDM-style -parameterization and reorder the noise levels such σ
that perceptual degradation progresses uniformly across
timesteps. The proposed method is designed to uniformly sample
and schedule values during both training and inference, σ
ensuring that the denoising capabilities learned across all noise
difficulty levels are consistently leveraged at inference time.
Furthermore, we incorporate trajectory refinement by reusing
reverse diffusion outputs for finetuning, which enhances
fine-grained color details and style alignment while preserving
structural stability. Experimental results demonstrate that
SSIMBaD achieves more balanced performance in terms of
structure, style, and realism compared to AnimeDiffusion, SCFT,
ControlNet, and train-free attention-based methods. Notably, our
approach yields significant improvements in FID and MS-SSIM
under cross-reference settings. Overall, this work shows that
SSIM-guided sigma reordering combined with reverse-trajectory
refinement constitutes a general and effective design principle for
conditional generation tasks where structural preservation is
critical.
목차 (Table of Contents)