Alattas, Eman
ORCID: https://orcid.org/0009-0004-8411-655X
(2025)
Deepfake Video Detection with CoAtNet Variations.
PhD thesis, University of Sheffield.
Abstract
Deepfakes are synthesised images, videos, audio or text generated using advanced artificial intelligence. They are often so realistic that humans struggle to determine their authenticity. Facial deepfakes pose significant societal risks, including disinformation, identity fraud, reputational damage, and biometric authentication attacks, making them one of the most important and widely studied forms of synthetic media. Despite extensive research, key challenges remain, most notably generalisation, computational cost, and temporal aggregation.
This thesis provides an extensive exploration of CoAtNet architecture variants for deepfake video detection, motivated by CoAtNet's ability to integrate convolutional inductive biases with attention-based global feature learning, examining generalisation and efficiency across frame- and video-level settings and in both within- and cross-dataset regimes. We first quantified the baseline CoAtNet performance using implicit features and then introduced explicit, landmark-driven representations to test whether facial patch cues can sustain performance while enabling more efficient models. The CoAtNet model processes single frames (images), but the results from multiple input frames can be aggregated. We also propose and evaluate a 3D-CoAtNet model that processes multiple frames simultaneously.
As contributions, the thesis provides a comprehensive CoAtNet framework that assesses generalisation for deepfake video detection, proposes an enhanced variant (CoAtNet16A) for improved cross-dataset performance, and systematically studies frame-selection strategies (middle, random, and optical-flow frames). We further introduce PatchCoAtNet, which replaces full-face inputs with images composed of landmark-based image ‘patches’ and uses ensemble learning to retain the detection performance at substantially lower input costs. Finally, we present 3D-CoAtNet, which inflates CoAtNet’s 2D convolutions, pooling, residual, and attention layers into a spatiotemporal (3D) architecture to provide an alternative to multiple frame results aggregation and to provide a means of modelling dependencies across consecutive frames for video-level detections. Collectively, these contributions advance the field by improving generalisation, efficiency and temporal modelling.
Metadata
| Supervisors: | Clark, John and Alsulami, Bassma |
|---|---|
| Keywords: | CoAtNet, Computational Efficiency, Computer Vision (CV), Convolutional Neural Networks (CNNs), Cross-Dataset Generalisation, Deepfake Detection, Digital Multimedia Forensics, Ensemble Learning, Facial Landmark Patches, Generalisation, Generative Adversarial Networks (GANs), Media Forensics, Resource-Efficient Models, Vision Transformers (ViTs) |
| Awarding institution: | University of Sheffield |
| Academic Units: | The University of Sheffield > Faculty of Engineering (Sheffield) > Computer Science (Sheffield) |
| Date Deposited: | 16 Sep 2026 11:05 |
| Last Modified: | 16 Sep 2026 11:05 |
| Open Archives Initiative ID (OAI ID): | oai:etheses.whiterose.ac.uk:39406 |
Download
Final eThesis - complete (pdf)
Embargoed until: 11 September 2027
This file cannot be downloaded or requested.
Filename: Alattas_Eman_DeepfakeVideoDetectionwithCoAtNetVariations.pdf
Export
Statistics
You can contact us about this thesis. If you need to make a general enquiry, please see the Contact us page.