For machine intelligence to extract high-level semantic information from video data and apply it to real-world interaction and problem-solving, two capabilities are essential: the ability to identify the entities present in a scene, and the ability to...
For machine intelligence to extract high-level semantic information from video data and apply it to real-world interaction and problem-solving, two capabilities are essential: the ability to identify the entities present in a scene, and the ability to recognize the dynamic interactions that unfold among them over time. Video Scene Graph Generation (VidSGG) provides a principled framework for such structured semantic extraction, representing a video sequence as a temporally evolving graph through the systematic parsing of ⟨subject, predicate, object⟩ triplets. In this thesis, we analyze the limitations of existing VidSGG approaches and propose a novel framework to address them.
Accurately modeling the evolution of inter-object relations as a video sequence progresses remains a fundamental challenge. This difficulty arises because the temporal aggregation strategies of existing methods lack sufficient mechanisms to capture the high variability of relational dynamics. Consequently, existing models suffer from Relational Semantic Inertia — an inability to promptly detect transition boundaries when inter-object relations shift.
To overcome this problem, we propose the Relation-Transition Aware Transformer (RTAformer), a framework that explicitly models changes in relation representations through three complementary modules and incorporates transition information during relation prediction. First, Spatio-Semantic Instance Alignment (SSIA) applies contrastive regularization to instance-pair representations across adjacent frames, forming temporally consistent representations for the same subject–object pair so that subsequent relation-query differences reflect genuine semantic transitions. Second, Disentangled Temporal Aggregation (DTA) separates the temporal refinement of instance and relation information into two independent streams: the instance stream performs persistence-oriented aggregation, while the relation stream is dedicated to transition-sensitive processing. Third, Transition-Aware Delta Injection (TADI) decomposes the adjacent-frame relation-query displacement into a direction component and a magnitude component, modulates the resulting transition signal, and injects it into the relation prediction query — enabling the model to respond promptly at relational transition boundaries.
The proposed method is validated on the Action Genome benchmark through quantitative and qualitative experiments, demonstrating superior performance over existing VidSGG methods. In particular, a dedicated evaluation on a constructed relation-transition subset shows that RTAformer improves transition-boundary prediction over representative existing methods, supporting the effectiveness of transition-aware temporal modeling in mitigating relational semantic inertia. The proposed framework is expected to advance machine intelligence by enabling more accurate semantic extraction from video data, with broad applicability to embodied AI, grounded reasoning, and holistic video understanding.