Thesis

Generative AI-driven multimedia semantic communication for wireless networks : semantic representation, robust transmission, and predictive modeling

Creator
Rights statement
Awarding institution
  • University of Strathclyde
Date of award
  • 2026
Thesis identifier
  • T18111
Person Identifier (Local)
  • 202258672
Qualification Level
Qualification Name
Department, School or Faculty
Abstract
  • The rapid integration of image and video-based intelligent applications into wireless networks has demanded communication paradigms that go beyond bit-level reconstruction. Semantic communication (SemCom) has emerged as a novel paradigm that shapes the next generation of wireless networks by transmitting “meaning”. However, existing designs suffer from critical gaps due to the lack of alignment with the evolving SemCom theories. These limitations are characterized by the absence of minimally sufficient semantic representations, limited robustness of learned semantics under channel noise, poor scalability across tasks and data distributions, and a lack of evaluation methodologies for semantic distortion, noise, and predictive sufficiency in image and video transmission. This research addresses these gaps through a theoretically grounded approach that leverages representation learning, discrete semantic abstraction, semantic coding, and generative artificial intelligence (AI) coupled with classical channel coding techniques. First, the thesis lays the foundation for minimally sufficient semantic symbol transmission for generative modeling by conveying semantic meaning as segmentation maps, achieving compression gains of up to 20× and stable semantic reconstruction in low signal-to-noise ratio (SNR) regions where conventional codecs fail. This study laid the groundwork for identifying theoretically grounded practical design considerations for subsequent research. Second, the study advances the concept of minimal semantic symbol transmission and explores vector-quantized and contrastively aligned semantic representations that function as shared semantic language between transmitter and receiver. This model is an advancement of edge-computing-based SemCom system initially developed, which has used a vector-quantized variational autoencoder (VQ-VAE) to ground representation learning into the semantic extraction process. By adopting index-based transmission protected by classical channel coding, mapped to a shared semantic codebook, the proposed systems achieve improved robustness, scalability, and multi-task accuracy without requiring channel-specific retraining. A Wasserstein-distance-based discriminator is introduced as the main novelty in this framework to stabilize adversarial training and guide learned representations toward a semantic manifold defined by the training image distribution under noisy transmission conditions. The model uses novel multilevel contrastive optimization at pixel, latent, perpetual, and semantic task levels to combat semantic and channel noise. In wireless image transmission, the model achieved a 56.2% reduction in symbol transmission load with lower model complexity and latency than joint source-channel coding (JSCC). In machine-perceived task performance, it achieved 94% classification accuracy, and human-perceived task performance, it achieved 91% image captioning accuracy for SNRs above 4 dB. The design demonstrated learned perceptual image patch similarity (LPIPS) of approximately 0.3 compared to 5.5 with better portable graphics (BPG) in fading channels. Third, the thesis presents a codebook optimization strategy and a formal quantification of semantic noise arising from symbol and representational ambiguities. A novel evaluation methodology based on the fast gradient sign method (FGSM) is introduced to evaluate the system’s robustness to semantic noise. The novel quantization methodology is presented using soft assignments, adversarial discrimination, and entropy-based regularization to minimize semantic ambiguity. The model achieved up to 96.88% codebook usage for codebook sizes of 64 or higher, with higher perplexity than vector quantized variational auto encoder (VQ-VAE). FGSM-based latent semantic noise evaluation showed classification accuracy of 60%-90% across the highest and lowest perturbation magnitudes.
Advisor / supervisor
  • Fernando, Anil
  • Dong, Feng
Resource Type
DOI

Relations

Items