Thesis
From contextual understanding to cognitive adaptation : multimodal AI for intelligent and adaptive multimedia systems
- Creator
- Rights statement
- Awarding institution
- University of Strathclyde
- Date of award
- 2026
- Thesis identifier
- T18117
- Person Identifier (Local)
- 202294952
- Qualification Level
- Qualification Name
- Department, School or Faculty
- Abstract
- The rapid evolution of streaming and on-demand media has transformed video platforms into adaptive, data-driven ecosystems in which content presentation, interruption timing, and user experience must be continuously optimised. Despite advances in multimodal machine learning, contemporary media decision systems remain largely perception-driven, operating on surface-level cues without structured reasoning over semantic meaning, narrative organisation, and viewer cognitive state. Addressing this challenge requires an integrated modular multimodal framework capable of jointly modelling content understanding, narrative structure, and human centred suitability. This thesis develops a modular, integrated multimodal reasoning architecture for adaptive media systems, enabling structured semantic grounding, narrative coherence modelling, and viewer-centred cognitive assessment, and applies this framework to context-aware transition and advertisement placement decisions. The first contribution establishes a multimodal topic–taxonomy mapping framework combining visual and transcript-derived features within a BERTopic pipeline for contextual advertising. The second contribution extends this architecture with semantically richer embeddings, replacing Bag-of-Visual-Words features with CLIP-ViT visual representations and incorporating image captioning and LLM-generated topic explanations to improve interpretability, multilingual robustness, and taxonomy alignment. Building on this foundation, the third contribution proposes a theme-aware segmentation framework that integrates deterministic perception modules with constrained, zero-temperature large language model reasoning. A controlled single pass Tree-of-Thoughts mechanism induces global narrative themes and identifies boundaries corresponding to sustained thematic transitions. Unlike stochastic or heavily fine-tuned alternatives, the method remains training-free, reproducible, and computationally lightweight while demonstrating improved boundary stability and higher agreement with human annotations across diverse video genres. The fourth contribution presents a multimodal cognitive-load estimation model that maps aligned audio–visual features onto NASA-TLX-inspired dimensions to generate a time-varying cognitive-suitability curve. Evaluation against human-annotated transition points shows that the model reliably identifies low-load narrative pauses and outperforms silence-based and rhythm-based heuristics, providing a complementary viewercentred segmentation signal. The fifth contribution synthesises these semantic, narrative, and cognitive components into an integrated modular segmentation and ad-break decision pipeline, operationalised through the AdBreakScore(t) framework. This integrated system jointly reasons about semantic coherence, thematic continuity, and cognitive readiness to produce context-aware boundary decisions under interpretability and deployment constraints. Together, this thesis advances adaptive multimedia systems from heuristic content analysis toward structured multimodal reasoning, establishing a human-centred framework applicable to context-aware advertising, narrative-aware editing assistance, and intelligent media adaptation.
- Advisor / supervisor
- Fernando, Anil
- Dong, Feng
- Resource Type
- DOI
Relations
Items
| Thumbnail | Title | Date Uploaded | Visibility | Actions |
|---|---|---|---|---|
|
|
PDF of thesis T18117 | 2026-08-25 | Public | Download |