RISS 학술연구정보서비스

검색

인기 검색어

    다국어 입력

    http://chineseinput.net/에서 pinyin(병음)방식으로 중국어를 변환할 수 있습니다.

    변환된 중국어를 복사하여 사용하시면 됩니다.

    예시)
    • 中文 을 입력하시려면 zhongwen을 입력하시고 space를누르시면됩니다.
    • 北京 을 입력하시려면 beijing을 입력하시고 space를 누르시면 됩니다.
    닫기

    Learning Structured Knowledge Representation for Video Understanding = 비디오 이해를 위한 구조적 지식 표현 학습

    한글로보기

    https://www.riss.kr/link?id=T17450553

    • 0

      상세조회
    • 0

      다운로드
    서지정보 열기
    • 내보내기
    • 내책장담기
    • 공유하기
    • 오류접수

    부가정보

    국문 초록 (Abstract) kakao i 다국어 번역

    인간의 지능은 세상을 구조적으로 인식하고, 그 관계를 탐구하기 위해 질문을 던지며, 이해를 바탕으로 목적 있는 행동을 계획하는 능력으로 특징지어진다. 본 논문은 이러한 인간의 인지 과정을 인공지능이 모방할 수 있도록, 비디오 이해를 위한 구조적 지식 표현(Structured Knowledge Representation, SKR)을 학습하고 활용하는 방법을 탐구한다. 본 연구의 궁극적인 목표는 저수준의 지각
    (perception)과 고수준의 인지적 추론(reasoning)을 연결하여, 복잡한 비디오 환경에서 구조적 지식을 기반으로 해석, 학습, 그리고 추론 할 수 있는 인공지능을 구축하는 것이다.

    이를 위해, 본 논문은 인간과 유사한 비디오 이해를 위한 통합 아키텍처를 제안하며, 이는 지각, 구조적 표현, 추론, 계획의 네 계층으로 구성된다. 이 아키텍처는 다중 모달 입력(프레임, 자막, 대사)을 구조적 의미 그래프로 변환하고, 이를 바탕으로 인지적 질문 생성을 통한 추론을 수행하며, 최종적으로 절차적 계획으로 확장되는 일련의 단계를 포함한다. 이러한 통합 구조는 본 논문의 세 가지 주요 연구 기여를 뒷받침한다. 첫째, 구조적 지식 표현의 생성 및 파싱을 위한 방법으로 엔티티 동의어 정렬 (Entity Synset Alignment, ESA)과 추상 의미 표현 기반 장면 그래프 파싱(Scene Graph Parsing via AMR, SGRAM)을 제안한다. 이는 시각적 · 언어적 정보를 통합된 의미 그래프 공간으로 정렬하여, 해석 가능하고 전이 가능한 구조적 표현 학습의 기반을 마련한다.

    둘째, INQUIRER라는 구조적 추론 프레임워크를 개발하여, 인공지능이 비디오에 대해 스스로 질문하고 대답하는 능력을 갖추도록 한다. INQUIRER는 다중 모달 단서를 기반으로 내부 지식 그래프(Internal Knowledge Graph, IKG)를 구축하고, 대규모 언어모델을 활용해 인지적으로 다양한 질문을 생성하며, 혼란도 (perplexity)를 이용한 질문 정제를 수행한다. 이를 통해 DramaQA, TVQA, STAR, How2QA 등 다양한 VideoQA 벤치마크를 구조적 QA 쌍으로 확장하여, 모델의 해석력과 성능을 동시에 향상시킨다.

    셋째, 이러한 구조적 추론을 절차적 계획(Procedure Planning)으로 접목/확장하기 위해 하이브리드 상태 표현(Hybrid State Representation, HSR)을 제안한다. HSR은 의미 상태 그래프(Semantic State Graph, SSG)와 문맥 기반 QA 쌍을 통합하여, 절차적 상태의 관계적 정확성과 문맥적 풍부함을 동시에 표현한다. 이를 통해 COIN, CrossTask, NIV 등의 벤치마크에서 상태 전이와 장기적 추론을 정밀하게 모델링하며, 기존 방법 대비 유의미한 성능 향상을 보인다.

    본 논문은 이러한 일련의 연구를 통해, 구조적 지식이 단순한 표현 도구를 넘어 인간 수준의 비디오 지능을 구현하기 위한 인지적 기반(cognitive substrate)임을 증명한다. 이는 인공지능이 복잡한 시각적 세계 속에서 관찰하고, 질문하며, 계획할 수 있는 새로운 패러다임을 제시한다.
    번역하기

    인간의 지능은 세상을 구조적으로 인식하고, 그 관계를 탐구하기 위해 질문을 던지며, 이해를 바탕으로 목적 있는 행동을 계획하는 능력으로 특징지어진다. 본 논문은 이러한 인간의 인지 ...

    인간의 지능은 세상을 구조적으로 인식하고, 그 관계를 탐구하기 위해 질문을 던지며, 이해를 바탕으로 목적 있는 행동을 계획하는 능력으로 특징지어진다. 본 논문은 이러한 인간의 인지 과정을 인공지능이 모방할 수 있도록, 비디오 이해를 위한 구조적 지식 표현(Structured Knowledge Representation, SKR)을 학습하고 활용하는 방법을 탐구한다. 본 연구의 궁극적인 목표는 저수준의 지각
    (perception)과 고수준의 인지적 추론(reasoning)을 연결하여, 복잡한 비디오 환경에서 구조적 지식을 기반으로 해석, 학습, 그리고 추론 할 수 있는 인공지능을 구축하는 것이다.

    이를 위해, 본 논문은 인간과 유사한 비디오 이해를 위한 통합 아키텍처를 제안하며, 이는 지각, 구조적 표현, 추론, 계획의 네 계층으로 구성된다. 이 아키텍처는 다중 모달 입력(프레임, 자막, 대사)을 구조적 의미 그래프로 변환하고, 이를 바탕으로 인지적 질문 생성을 통한 추론을 수행하며, 최종적으로 절차적 계획으로 확장되는 일련의 단계를 포함한다. 이러한 통합 구조는 본 논문의 세 가지 주요 연구 기여를 뒷받침한다. 첫째, 구조적 지식 표현의 생성 및 파싱을 위한 방법으로 엔티티 동의어 정렬 (Entity Synset Alignment, ESA)과 추상 의미 표현 기반 장면 그래프 파싱(Scene Graph Parsing via AMR, SGRAM)을 제안한다. 이는 시각적 · 언어적 정보를 통합된 의미 그래프 공간으로 정렬하여, 해석 가능하고 전이 가능한 구조적 표현 학습의 기반을 마련한다.

    둘째, INQUIRER라는 구조적 추론 프레임워크를 개발하여, 인공지능이 비디오에 대해 스스로 질문하고 대답하는 능력을 갖추도록 한다. INQUIRER는 다중 모달 단서를 기반으로 내부 지식 그래프(Internal Knowledge Graph, IKG)를 구축하고, 대규모 언어모델을 활용해 인지적으로 다양한 질문을 생성하며, 혼란도 (perplexity)를 이용한 질문 정제를 수행한다. 이를 통해 DramaQA, TVQA, STAR, How2QA 등 다양한 VideoQA 벤치마크를 구조적 QA 쌍으로 확장하여, 모델의 해석력과 성능을 동시에 향상시킨다.

    셋째, 이러한 구조적 추론을 절차적 계획(Procedure Planning)으로 접목/확장하기 위해 하이브리드 상태 표현(Hybrid State Representation, HSR)을 제안한다. HSR은 의미 상태 그래프(Semantic State Graph, SSG)와 문맥 기반 QA 쌍을 통합하여, 절차적 상태의 관계적 정확성과 문맥적 풍부함을 동시에 표현한다. 이를 통해 COIN, CrossTask, NIV 등의 벤치마크에서 상태 전이와 장기적 추론을 정밀하게 모델링하며, 기존 방법 대비 유의미한 성능 향상을 보인다.

    본 논문은 이러한 일련의 연구를 통해, 구조적 지식이 단순한 표현 도구를 넘어 인간 수준의 비디오 지능을 구현하기 위한 인지적 기반(cognitive substrate)임을 증명한다. 이는 인공지능이 복잡한 시각적 세계 속에서 관찰하고, 질문하며, 계획할 수 있는 새로운 패러다임을 제시한다.

    더보기

    다국어 초록 (Multilingual Abstract) kakao i 다국어 번역

    Human intelligence is characterized by its ability to perceive structured relationships in the world, ask questions to reason about them, and plan purposeful actions based on inferred understanding. This dissertation explores how artificial intelligence can emulate such human-like cognition through the construction and utilization of Structured Knowledge Representation (SKR) for video understanding. The overarching goal is to bridge low-level perception and high-level cognitive reasoning by enabling AI systems to parse, learn/reason, and act over structured knowledge derived from complex video environments.

    To this end, I propose Perception-Reasoning-Interaction for Structured Knowledge Modeling (PRISM), a unified architecture for human-like video understanding that spans three interconnected layers: structured representation, reasoning, and planning. The architecture begins by transforming multimodal video inputs—frames, captions, and subtitles—into structured semantic graphs, followed by reasoning over these representations through cognitive inquiry, and ultimately extends to procedural planning that mirrors human decision-making. This architecture forms the backbone for the three major contributions of this dissertation.

    First, I introduce methods for structured knowledge representation generation, including Entity Synset Alignment (ESA) and Scene Graph Parsing via Abstract Meaning Representation (SGRAM). These approaches unify heterogeneous visual and linguistic signals into a coherent semantic graph space, establishing the foundation for interpretable and transferable structured representations.

    Building upon this foundation, I develop INQUIRER, a structured reasoning framework that enables AI systems to ask and answer questions about videos. INQUIRER constructs Internal Knowledge Graphs (IKGs) from multimodal cues, generates cognitively diverse questions through large language models, and filters them using perplexity-based curation. By augmenting existing VideoQA benchmarks (DramaQA, TVQA, STAR, and How2QA) with QA pairs, INQUIRER enhances both the interpretability and performance of video reasoning models, bridging human-like curiosity and machine understanding.

    Finally, I extend this structured reasoning paradigm to procedure planning through a Hybrid State Representation (HSR). HSR integrates Semantic State Graphs (SSGs) with contextual QA pairs, capturing both relational precision and contextual clarity of procedural states. This hybrid representation enables accurate modeling of state transitions and long-horizon reasoning, leading to substantial performance improvements across COIN, CrossTask, and NIV benchmarks.

    Together, these contributions present a coherent cognitive trajectory from structured perception to reasoning and planning. The dissertation demonstrates that structured knowledge is not merely a descriptive medium for visual understanding, but a cognitive substrate—a foundation upon which human-like video intelligence can emerge, enabling AI systems to observe, question, and plan within complex visual worlds.
    번역하기

    Human intelligence is characterized by its ability to perceive structured relationships in the world, ask questions to reason about them, and plan purposeful actions based on inferred understanding. This dissertation explores how artificial intelligen...

    Human intelligence is characterized by its ability to perceive structured relationships in the world, ask questions to reason about them, and plan purposeful actions based on inferred understanding. This dissertation explores how artificial intelligence can emulate such human-like cognition through the construction and utilization of Structured Knowledge Representation (SKR) for video understanding. The overarching goal is to bridge low-level perception and high-level cognitive reasoning by enabling AI systems to parse, learn/reason, and act over structured knowledge derived from complex video environments.

    To this end, I propose Perception-Reasoning-Interaction for Structured Knowledge Modeling (PRISM), a unified architecture for human-like video understanding that spans three interconnected layers: structured representation, reasoning, and planning. The architecture begins by transforming multimodal video inputs—frames, captions, and subtitles—into structured semantic graphs, followed by reasoning over these representations through cognitive inquiry, and ultimately extends to procedural planning that mirrors human decision-making. This architecture forms the backbone for the three major contributions of this dissertation.

    First, I introduce methods for structured knowledge representation generation, including Entity Synset Alignment (ESA) and Scene Graph Parsing via Abstract Meaning Representation (SGRAM). These approaches unify heterogeneous visual and linguistic signals into a coherent semantic graph space, establishing the foundation for interpretable and transferable structured representations.

    Building upon this foundation, I develop INQUIRER, a structured reasoning framework that enables AI systems to ask and answer questions about videos. INQUIRER constructs Internal Knowledge Graphs (IKGs) from multimodal cues, generates cognitively diverse questions through large language models, and filters them using perplexity-based curation. By augmenting existing VideoQA benchmarks (DramaQA, TVQA, STAR, and How2QA) with QA pairs, INQUIRER enhances both the interpretability and performance of video reasoning models, bridging human-like curiosity and machine understanding.

    Finally, I extend this structured reasoning paradigm to procedure planning through a Hybrid State Representation (HSR). HSR integrates Semantic State Graphs (SSGs) with contextual QA pairs, capturing both relational precision and contextual clarity of procedural states. This hybrid representation enables accurate modeling of state transitions and long-horizon reasoning, leading to substantial performance improvements across COIN, CrossTask, and NIV benchmarks.

    Together, these contributions present a coherent cognitive trajectory from structured perception to reasoning and planning. The dissertation demonstrates that structured knowledge is not merely a descriptive medium for visual understanding, but a cognitive substrate—a foundation upon which human-like video intelligence can emerge, enabling AI systems to observe, question, and plan within complex visual worlds.

    더보기

    목차 (Table of Contents)

    • Abstract i
    • Chapter 1 Introduction 1
    • 1.1 Approaches and Contributions 1
    • 1.2 Thesis Organization 2
    • Chapter 2 Background 4
    • Abstract i
    • Chapter 1 Introduction 1
    • 1.1 Approaches and Contributions 1
    • 1.2 Thesis Organization 2
    • Chapter 2 Background 4
    • 2.1 Structured Knowledge Representation 4
    • 2.1.1 Scene Graph 4
    • 2.1.2 Abstract Meaning Representation 5
    • 2.1.3 Knowledge Graph 6
    • 2.2 Reasoning with Structured Knowledge Representation 6
    • 2.2.1 Visual Question Answering with Structured Knowledge 6
    • 2.2.2 Video Question Answering 7
    • 2.2.3 Video Question Generation 8
    • 2.2.4 Video Procedure Planning 9
    • Chapter 3 Video Understanding Architecture 11
    • 3.1 Overview 11
    • 3.2 System Architecture and Information Flow 13
    • Chapter 4 Structured Knowledge Representation Generation 15
    • 4.1 Introduction 15
    • 4.2 Method 18
    • 4.2.1 Entity Synset Alignment for General Scene Graph Construction 18
    • 4.2.2 AMR-based Scene Graph Parsing 21
    • 4.3 Experimental Setup 25
    • 4.3.1 Datasets 25
    • 4.3.2 Implementation Details 26
    • 4.4 Results 27
    • 4.4.1 Quantitative Results 27
    • 4.4.2 Qualitative Results 31
    • 4.4.3 Ablation Study 36
    • 4.5 Discussion and Summary 39
    • Chapter 5 Structured Knowledge-Guided Video Reasoning 41
    • 5.1 Introduction 41
    • 5.2 Framework Overview 44
    • 5.3 Methodology 46
    • 5.3.1 Knowledge Construction Module 47
    • 5.3.2 Question Generation Module 49
    • 5.3.3 Question Curation Module 54
    • 5.4 Dataset Augmentation on VideoQA benchmarks 55
    • 5.5 Experimental Setup 58
    • 5.5.1 Baselines and Implementation Details 58
    • 5.5.2 Dataset 59
    • 5.6 Results and Analysis 60
    • 5.6.1 Quality for generated QA pairs 61
    • 5.6.2 Effectiveness of IKGs 61
    • 5.6.3 Human evaluation 63
    • 5.6.4 Ablation Study 65
    • 5.6.5 Qualitative Analysis 68
    • 5.7 Discussion and Summary 69
    • Chapter 6 Structured Knowledge-Guided Video Procedure Plan-
    • ning 73
    • 6.1 Introduction 73
    • 6.2 Method 76
    • 6.2.1 Problem formulation 76
    • 6.2.2 Hybrid State Representation 77
    • 6.2.3 Heterogeneous State Encoder 77
    • 6.2.4 Action Step Decoding 78
    • 6.2.5 Learning Objective and Inference 79
    • 6.3 Experimental Setup 80
    • 6.3.1 Datasets 80
    • 6.3.2 Evaluation metrics 80
    • 6.3.3 Baselines 81
    • 6.3.4 Implementation details 81
    • 6.4 Results 82
    • 6.4.1 Quantitative Results 82
    • 6.4.2 Ablation Study 85
    • 6.4.3 Qualitative Results 87
    • 6.5 Discussion and Summary 89
    • Chapter 7 Concluding Remarks 92
    • 7.1 Summary 92
    • 7.2 Future Work 94
    • Appendix A Appendix 97
    • A.1 Appendix for Video Reasoning 97
    • A.1.1 Distribution of question types in DramaQA 97
    • A.1.2 Qualitative Examples of INQUIRER vs. Naive 99
    • A.1.3 Samples of knowledge structure and QGen prompt 99
    • A.1.4 Human evaluation questionnaire 100
    • A.2 Appendix for Video Procedure Planning 100
    • A.2.1 Datasets 100
    • A.2.2 Prompt Design 102
    • A.2.3 Validity and Robustness of Semantic State Graph 102
    • A.2.4 Detailed Formulation of Visual-State Alignment 104
    • A.2.5 Extended Qualitative Analysis 105
    • Acknowledgements 126
    • 요약 128
    더보기

    분석정보

    View

    상세정보조회

    0

    Usage

    원문다운로드

    0

    대출신청

    0

    복사신청

    0

    EDDS신청

    0

    동일 주제 내 활용도 TOP

    더보기

    주제

    연도별 연구동향

    연도별 활용동향

    연관논문

    연구자 네트워크맵

    공동연구자 (7)

    유사연구자 (20) 활용도상위20명

    이 자료와 함께 이용한 RISS 자료

    나만을 위한 추천자료

    해외이동버튼