RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Evaluating Multi-Scene Video Generation for Long-Form Content = 장편 콘텐츠를 위한 다중장면 비디오 생성 평가

    한글로보기

    https://www.riss.kr/link?id=T17450272

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    최근 Diffusion 모델의 발전으로 실용적 수준의 고품질 비디오 생성이 가능 합니다 . 그러나 여러 장면이 연결된 영화 수준의 콘텐츠를 제작하는 것은 여전히 어려운 과제입니다 . 핵심 난제는 서사적 일관성과 시각적 일관성을 유지하는 것입니다 . 기존 방법들은 캐릭터와 배경 간 혼합 및 왜곡과 같은 문제를 겪는데 , 주로 제어되지 않은 확률적 샘플링에 기인합니다 . 이로 인해 사용자들은 수많은 후보를 생성해야 하며 , 자동화를 위한 통합적이고 적응적인 평가 점수가 없어 모든 후보를 검증하는 과정이 매우 노동 집약적인 병목 지점이 됩니다 . 또 다른 중요한 문제는 평가 품질과 실행 시간 성능 간의 상충관계입니다 . 인간과 유사한 판단을 가장 잘 포착하는 지표들은 종종 반복적 생성을 지원하기에는 너무 느립니다 . 효과적인 평가의 부재에서 비롯된 이러한 문 제들이 연구의 새로운 솔루션 개발 동기가 되었습니다 . 이를 해결하기 위해 세 가지 핵 심 솔루션을 제안합니다 . 첫째 , Attention 구조를 사용하여 생성된 비디오를 적응적으로 평가하는 MSG(Multi Scene Generation) 점수를 도입합니다. 둘째 , 다수의 후보 중에서 최상의 결과를 자동 식별하는 CGS(Candidate Generation and Selection) 구조를 제시합니다 . 마지막으로 , IID(Implicit Insight Distillation) 방법은 평가의 품질과 속도 간의 상충관계를 해결합니다 . 이는 전체 지표 모음에 대 한 교사 모델의 통찰을 학생 모델로 증류하여 , 높은 정확도를 유지하면서 고속 평가를 가능하게 합니다 . 이러한 기여는 장편 비디오 생성을 위한 포괄적 솔루션을 제공합니다.
    번역하기

    최근 Diffusion 모델의 발전으로 실용적 수준의 고품질 비디오 생성이 가능 합니다 . 그러나 여러 장면이 연결된 영화 수준의 콘텐츠를 제작하는 것은 여전히 어려운 과제입니다 . 핵심 난제는 ...

    최근 Diffusion 모델의 발전으로 실용적 수준의 고품질 비디오 생성이 가능 합니다 . 그러나 여러 장면이 연결된 영화 수준의 콘텐츠를 제작하는 것은 여전히 어려운 과제입니다 . 핵심 난제는 서사적 일관성과 시각적 일관성을 유지하는 것입니다 . 기존 방법들은 캐릭터와 배경 간 혼합 및 왜곡과 같은 문제를 겪는데 , 주로 제어되지 않은 확률적 샘플링에 기인합니다 . 이로 인해 사용자들은 수많은 후보를 생성해야 하며 , 자동화를 위한 통합적이고 적응적인 평가 점수가 없어 모든 후보를 검증하는 과정이 매우 노동 집약적인 병목 지점이 됩니다 . 또 다른 중요한 문제는 평가 품질과 실행 시간 성능 간의 상충관계입니다 . 인간과 유사한 판단을 가장 잘 포착하는 지표들은 종종 반복적 생성을 지원하기에는 너무 느립니다 . 효과적인 평가의 부재에서 비롯된 이러한 문 제들이 연구의 새로운 솔루션 개발 동기가 되었습니다 . 이를 해결하기 위해 세 가지 핵 심 솔루션을 제안합니다 . 첫째 , Attention 구조를 사용하여 생성된 비디오를 적응적으로 평가하는 MSG(Multi Scene Generation) 점수를 도입합니다. 둘째 , 다수의 후보 중에서 최상의 결과를 자동 식별하는 CGS(Candidate Generation and Selection) 구조를 제시합니다 . 마지막으로 , IID(Implicit Insight Distillation) 방법은 평가의 품질과 속도 간의 상충관계를 해결합니다 . 이는 전체 지표 모음에 대 한 교사 모델의 통찰을 학생 모델로 증류하여 , 높은 정확도를 유지하면서 고속 평가를 가능하게 합니다 . 이러한 기여는 장편 비디오 생성을 위한 포괄적 솔루션을 제공합니다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Recent advances in text-to-video diffusion models have enabled the generation of high-quality video at a practical level. However, creating film-level content composed of multiple interconnected scenes remains a challenge.
    The core difficulty lies in maintaining narrative coherence and visual consistency. Existing methods often suffer from artifacts such as character-background blending and distortion, largely due to uncontrolled probabilistic sampling. This requires generating numerous candidates, and the process of verifying them all is a highly burden bottleneck due to the lack of a unified, adaptive evaluation score for automation.
    Another critical issue is the trade-off between evaluation quality and run-time performance: metrics that best capture human-like judgment are often too slow to support iterative generation. These challenges, originating from the lack of an effective evaluation, motivate our work toward a novel solution.
    To address these challenges, we propose three core solutions. First, we introduce the MSG(Multi-Scene Generation) score, which uses a hierarchical temporal and spatial attention structure to adaptively evaluate generated videos. Second, we present the CGS(Candidate Generation and Selection) framework that automatically identifies the best result among many candidates. Finally, the IID(Implicit Insight Distillation) method resolves the trade-off between quality and speed of evaluation. It distills the insights from a teacher model that reasons over full suites of metrics into an efficient student, enabling high-speed evaluation while preserving a high accuracy. These contributions offer the film-level generation solution at a practical workflow.
    번역하기

    Recent advances in text-to-video diffusion models have enabled the generation of high-quality video at a practical level. However, creating film-level content composed of multiple interconnected scenes remains a challenge. The core difficulty lies in ...

    Recent advances in text-to-video diffusion models have enabled the generation of high-quality video at a practical level. However, creating film-level content composed of multiple interconnected scenes remains a challenge.
    The core difficulty lies in maintaining narrative coherence and visual consistency. Existing methods often suffer from artifacts such as character-background blending and distortion, largely due to uncontrolled probabilistic sampling. This requires generating numerous candidates, and the process of verifying them all is a highly burden bottleneck due to the lack of a unified, adaptive evaluation score for automation.
    Another critical issue is the trade-off between evaluation quality and run-time performance: metrics that best capture human-like judgment are often too slow to support iterative generation. These challenges, originating from the lack of an effective evaluation, motivate our work toward a novel solution.
    To address these challenges, we propose three core solutions. First, we introduce the MSG(Multi-Scene Generation) score, which uses a hierarchical temporal and spatial attention structure to adaptively evaluate generated videos. Second, we present the CGS(Candidate Generation and Selection) framework that automatically identifies the best result among many candidates. Finally, the IID(Implicit Insight Distillation) method resolves the trade-off between quality and speed of evaluation. It distills the insights from a teacher model that reasons over full suites of metrics into an efficient student, enabling high-speed evaluation while preserving a high accuracy. These contributions offer the film-level generation solution at a practical workflow.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Contents ii
    • List of Figures iv
    • List of Tables v
    • 1 Introduction 1
    • Abstract i
    • Contents ii
    • List of Figures iv
    • List of Tables v
    • 1 Introduction 1
    • 2 Related Works 4
    • 2.1 Video Generation 4
    • 2.2 Video Generation Evaluation 5
    • 3 Method 7
    • 3.1 MSG(Multi scene Generation) Score 7
    • 3.2 CGS(Candidate Generation and Selection) Framework 11
    • 3.3 IID(Implicit Insight Distillation) method 13
    • 4 Experiments 15
    • 4.1 Efficiency Compared to Manual Methods 15
    • 4.2 Robustness of the MSG Score 16
    • 4.3 Validating Against Preliminary Approaches 21
    • 4.4 Limitations 24
    • 5 Conclusion 25
    • 6 Supplementary Materials 27
    • 6.1 Detailed Explanation of Evaluation Submetrics 27
    • 6.2 Video Generation Samples And MSG Ranking 30
    • 6.2.1 "The Ant and the Grasshopper" 30
    • 6.2.2 "Aliens in Red Planet" 32
    • Bibliography 36
    • 초록 39
    • 감사의 글 41
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼