프레젠테이션을 위한 비디오는 과거 카메라 영상을 기반으로 많은 수작업을 통해 편집 및 제작되었지만, 현재 문서 내용을 자동으로 요약하는 인공지능, 비디오 슬라이드 영상을 생성해 주...

http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.
변환된 중국어를 복사하여 사용하시면 됩니다.
프레젠테이션을 위한 비디오는 과거 카메라 영상을 기반으로 많은 수작업을 통해 편집 및 제작되었지만, 현재 문서 내용을 자동으로 요약하는 인공지능, 비디오 슬라이드 영상을 생성해 주...
프레젠테이션을 위한 비디오는 과거 카메라 영상을 기반으로 많은 수작업을 통해 편집 및 제작되었지만, 현재 문서 내용을 자동으로 요약하는 인공지능, 비디오 슬라이드 영상을 생성해 주는 인공지능 등 다양한 인공지능 기술이 발전하고 있다. 본 연구에서는 이러한 다양한 인공지능을 사용하여 통합 스크립트(Integrated Script for Presentation Control)를 기반으로 한 클론 AI 프레젠테이션 비디오를 자동으로 생성하는 시스템을 제안하고, 이를 시스템으로 구현하고 성능을 평가하였다. 통합 스크립트는 사용자의 파워포인트 슬라이드의 텍스트를 추출하여 요약 문장을 만들고 사전에 학습된 사용자의 음성과 얼굴 영상을 가지는 클론 AI의 합성된 음성과 영상을 생성하는 스크립트로 구성되어 있다. 통합 스크립트에는 각 프레젠테이션 발표 자료의 내용 중 연출 효과가 필요한 구간에 마크업을 포함해 비디오 생성 시 텍스트에 특수한 효과를 나타내도록 설계하였다. 또한, 스크립트에는 발표자의 합성된 비디오와 동기화되어 텍스트 효과뿐만 아니라 지정된 시점에 동기화된 비디오가 생성되도록 하였다. 따라서 제안된 스크립트를 통해 사용자는 별도의 편집 과정 없이 각 슬라이드의 시각적 및 청각적 요소를 갖춘 프레젠테이션 자료와 시간상으로 동기화된 클론 AI의 프레젠테이션 비디오를 생성할 수 있다. 제안된 시스템은 음성 합성모듈, 얼굴 합성모듈 그리고 통합 스크립트에 의한 비디오 생성 모듈이 파이프라인으로 구성하였으며, 각 모듈은 독립적으로 설계되어 더 나은 성능의 인공지능 모듈로 대체될 수 있도록 하였다. 특히, 음성합성 모듈에서는 Fast-HuBERT 기반 오디오 특징 추출과정을 ONNX 추론 최적화 기법을 적용하여 연산 처리에 효율 향상을 가져왔으며, 이는 고성능 GPU 환경뿐만 아니라 모바일 및 임베디드와 같은 온디바이스 환경에서도 동작할 수 있음을 확인하였다. 따라서 본 논문에서 제안한 통합 스크립트 기반 프레젠테이션 비디오 생성 시스템은 클론 AI의 음성과 영상 그리고 통합 스크립트에 의한 프레젠테이션 슬라이드에 텍스트 효과가 설정된 시점과 동기화되어 나타나는 프레젠테이션 비디오를 자동으로 생성할 수 있다. 앞으로 제안된 통합 스크립트를 사용하여 다양한 효과로 표현된 자기만의 비디오를 생성할 수 있는 효율적인 방안이 마련되었다.
다국어 초록 (Multilingual Abstract)
Advances in artificial intelligence have transformed presentation creation from a manual process to an automated pipeline capable of summarizing documents and generating slides. This study extends this trend by proposing a personalized presentation vi...
Advances in artificial intelligence have transformed presentation creation from a manual process to an automated pipeline capable of summarizing documents and generating slides. This study extends this trend by proposing a personalized presentation video generation system that leverages an Integrated Script for Presentation Control (ISPC). The system was implemented and its performance evaluated. ISPC generates synthetic speech and video based on a user-written Script and a Clone AI pre-trained on user-specific audio-visual data. The Script contains markup indicating segments that require visual effects, and slide text effects are synchronized with the synthetic presenter video to the specified timestamps, producing a complete video. This allows users to create videos in which the visual and auditory components of each slide are temporally aligned without additional editing. The system architecture connects speech synthesis models, face synthesis models, ISPC, and video generation processes in a pipeline, with each module independently designed and replaceable. Additionally, the speech synthesis model applies Fast-HuBERT-based audio feature extraction and ONNX inference optimization to improve computational efficiency, and the system supports deployment on both high-performance GPU environments and on-device environments such as mobile and embedded platforms. The proposed ISPC-based system enables users to generate personalized presentation videos by synchronizing synthetic speech, video and slide text effects according to the timestamps defined in the Script. This approach allows users to regenerate videos in accordance with the Script settings without the need for manual editing. An efficient approach is now available to create unique videos with various effects by utilizing the proposed integrated script.
목차 (Table of Contents)