Abstract
Multimodal video understanding is evolving from short-clip visual recognition toward complex system-level intelligence that integrates heterogeneous modalities, long-horizon temporal structures, high-level semantic reasoning, cross-modal generation, and real-time interaction. As an immersive medium, video contains not only visual frames, audio signals, and textual cues, but also interaction-related information such as controller inputs, and other sensor streams in embodied environments. This expansion has changed the research scope of video understanding: models are no longer expected merely to recognize visible events or generate surface-level descriptions, but to align multimodal information, infer deeper semantic relations, support affective content generation, and respond interactively in real time.Despite recent progress in video Transformers, multimodal pretraining, and large multimodal video models, existing methods still face several limitations in complex real-world scenarios. First, many models remain insufficient for long-horizon multimodal semantic understanding in domain-specific contexts, where videos often contain extended temporal dependencies, heterogeneous input sources, and high-level domain semantics. Second, current research has not fully addressed how video understanding results can be transformed into stable and controllable constraints for cross-modal generation, particularly in affective scenarios such as video background music generation. Third, conventional video understanding systems are mostly designed for offline analysis and remain inadequate for real-time embodied interaction, where systems must continuously perceive multimodal user inputs, infer interaction intent, and generate coordinated responses through dialogue, facial expression, lip synchronization, and body motion.
To address these challenges, this thesis systematically investigates complex multimodal video understanding from three complementary perspectives: domain-specific semantic reasoning, affect-aware cross-modal generation, and real-time embodied interaction.
First, this thesis proposes a general multimodal video understanding framework for long-horizon, multi-source, and high-semantics video content in specialized domains. The framework incorporates segment-level recurrent semantic accumulation, multimodal parallel encoding and cross-semantic alignment, external knowledge integration, retrieval-augmented reasoning, domain-specific dataset construction, and progressive training. These designs aim to improve the model’s ability to represent extended video context, align heterogeneous modalities, incorporate domain knowledge, and generate fine-grained semantic interpretations.
Second, this thesis develops an emotion-aware video background music generation method that transforms multimodal video understanding results into controllable musical structures. The proposed method models the relationships among video semantics, narrative progression, affective trajectory, textual prompts, and musical structure. It introduces an emotion-aware Transformer architecture with text-interactive control, a decoupled mapping-generation design, and an end-to-end recurrent generation mechanism for improving long-range musical coherence. In addition, the thesis constructs the MovGM dataset from movie clips and cinematic soundtracks, and proposes a human-centered repainting test to evaluate whether generated music is perceptually appropriate when reinserted into the corresponding video.
Third, this thesis constructs a real-time multimodal interactive digital human system for virtual reality scenarios. The system extends multimodal video understanding from offline semantic analysis to an interactive embodied intelligence framework. It supports a continuous perception-understanding-reasoning-feedback loop by integrating multi-source signals such as speech, gestures, motion, viewpoint, and controller inputs. To generate natural and coordinated responses, the system introduces temporal-semantic joint alignment, adaptive multimodal fusion, explicit-implicit partitioned understanding, dual-center switchable reasoning, dual-source animation generation, and multi-level adaptation of character, content, voice, and motion. It further develops a multi-source data framework and benchmark suite for evaluating module-level performance, end-to-end interaction quality, and user-centered experience.
Overall, this thesis demonstrates that complex multimodal video understanding should not be treated as a single task of passive video analysis. Instead, it should be understood as a foundational capability that supports semantic reasoning, affective cross-modal generation, and real-time embodied interaction. By organizing these three studies under a unified research trajectory, the thesis provides methodological foundations, data resources, system designs, and evaluation protocols for advancing complex multimodal video understanding toward reasoning-oriented, generation-oriented, and interaction-oriented real-world applications.
| Date of Award | 15 Nov 2026 |
|---|---|
| Original language | English |
| Awarding Institution |
|
| Supervisor | Boon Giin Lee (Supervisor), Wooi Ping Cheah (Supervisor) & Yongmei Cai (Supervisor) |
Cite this
- Standard