최근 Diffusion 모델의 발전으로 실용적 수준의 고품질 비디오 생성이 가능 합니다 . 그러나 여러 장면이 연결된 영화 수준의 콘텐츠를 제작하는 것은 여전히 어려운 과제입니다 . 핵심 난제는 ...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
최근 Diffusion 모델의 발전으로 실용적 수준의 고품질 비디오 생성이 가능 합니다 . 그러나 여러 장면이 연결된 영화 수준의 콘텐츠를 제작하는 것은 여전히 어려운 과제입니다 . 핵심 난제는 ...
최근 Diffusion 모델의 발전으로 실용적 수준의 고품질 비디오 생성이 가능 합니다 . 그러나 여러 장면이 연결된 영화 수준의 콘텐츠를 제작하는 것은 여전히 어려운 과제입니다 . 핵심 난제는 서사적 일관성과 시각적 일관성을 유지하는 것입니다 . 기존 방법들은 캐릭터와 배경 간 혼합 및 왜곡과 같은 문제를 겪는데 , 주로 제어되지 않은 확률적 샘플링에 기인합니다 . 이로 인해 사용자들은 수많은 후보를 생성해야 하며 , 자동화를 위한 통합적이고 적응적인 평가 점수가 없어 모든 후보를 검증하는 과정이 매우 노동 집약적인 병목 지점이 됩니다 . 또 다른 중요한 문제는 평가 품질과 실행 시간 성능 간의 상충관계입니다 . 인간과 유사한 판단을 가장 잘 포착하는 지표들은 종종 반복적 생성을 지원하기에는 너무 느립니다 . 효과적인 평가의 부재에서 비롯된 이러한 문 제들이 연구의 새로운 솔루션 개발 동기가 되었습니다 . 이를 해결하기 위해 세 가지 핵 심 솔루션을 제안합니다 . 첫째 , Attention 구조를 사용하여 생성된 비디오를 적응적으로 평가하는 MSG(Multi Scene Generation) 점수를 도입합니다. 둘째 , 다수의 후보 중에서 최상의 결과를 자동 식별하는 CGS(Candidate Generation and Selection) 구조를 제시합니다 . 마지막으로 , IID(Implicit Insight Distillation) 방법은 평가의 품질과 속도 간의 상충관계를 해결합니다 . 이는 전체 지표 모음에 대 한 교사 모델의 통찰을 학생 모델로 증류하여 , 높은 정확도를 유지하면서 고속 평가를 가능하게 합니다 . 이러한 기여는 장편 비디오 생성을 위한 포괄적 솔루션을 제공합니다.
다국어 초록 (Multilingual Abstract)
Recent advances in text-to-video diffusion models have enabled the generation of high-quality video at a practical level. However, creating film-level content composed of multiple interconnected scenes remains a challenge. The core difficulty lies in ...
Recent advances in text-to-video diffusion models have enabled the generation of high-quality video at a practical level. However, creating film-level content composed of multiple interconnected scenes remains a challenge.
The core difficulty lies in maintaining narrative coherence and visual consistency. Existing methods often suffer from artifacts such as character-background blending and distortion, largely due to uncontrolled probabilistic sampling. This requires generating numerous candidates, and the process of verifying them all is a highly burden bottleneck due to the lack of a unified, adaptive evaluation score for automation.
Another critical issue is the trade-off between evaluation quality and run-time performance: metrics that best capture human-like judgment are often too slow to support iterative generation. These challenges, originating from the lack of an effective evaluation, motivate our work toward a novel solution.
To address these challenges, we propose three core solutions. First, we introduce the MSG(Multi-Scene Generation) score, which uses a hierarchical temporal and spatial attention structure to adaptively evaluate generated videos. Second, we present the CGS(Candidate Generation and Selection) framework that automatically identifies the best result among many candidates. Finally, the IID(Implicit Insight Distillation) method resolves the trade-off between quality and speed of evaluation. It distills the insights from a teacher model that reasons over full suites of metrics into an efficient student, enabling high-speed evaluation while preserving a high accuracy. These contributions offer the film-level generation solution at a practical workflow.
목차 (Table of Contents)