Clarke, Jason
ORCID: 0000-0002-1032-6472
(2026)
Robust Audiovisual Active Speaker Detection for Egocentric Recordings - Leveraging Contextual Information and Cross-Modal Biometric Associations.
PhD thesis, University of Sheffield.
Abstract
The task of Audiovisual Active Speaker Detection (ASD) involves determining when a visible person in a video is speaking by jointly analysing audiovisual cues indicative of speech. Conventional ASD systems heavily rely on modelling the cross-modal synchronicity between audio-visual cues that arise during periods of active speech. This involves modelling finegrained articulatory motion in the video signal and detecting speech in the audio signal. In egocentric recordings captured from wearable devices, rapid camera motion, partial occlusions, ambient noise, and computational resource constraints degrade the effectiveness of conventional synchronisation-based models.
This thesis explores how contextual and cross-modal biometric information can compensate for these limitations, advancing lightweight ASD suitable for real-world, on-device operation. Rather than relying solely on frame-level alignment between facial motion and speech, three complementary strategies are developed that progressively expand the contextual scope of ASD modelling.
First, full-scene embeddings derived from transformer-based architectures are introduced to capture the global contextual information provided by full-scene images, allowing baseline ASD models to exploit priors present in the full-scene semantics. Second, speaker-aware conditioning is achieved by proposing a system which extracts and compares vocal-identity embeddings, enabling disambiguation of visually degraded scenarios. Finally, a synchronisation-free framework based on cross-modal Face-Voice Association (FVA) is proposed, where shared biometric embeddings of faces and voices attribute speech to the relevant visible identity implicitly.
Together, these contributions demonstrate that contextual reasoning and cross-modal
biometric correspondence can substantially improve ASD robustness under egocentric conditions while preserving parameter efficiency. The resulting models achieve competitive or superior performance to heavier synchronisation-based systems on large-scale benchmarks. More broadly, the findings highlight the potential of context-enriched multimodal representations as a foundation for future wearable and Augmented Reality (AR)-based audiovisual diarisation systems.
Metadata
| Supervisors: | Goetze, Stefan and Gotoh, Yoshihiko |
|---|---|
| Related URLs: |
|
| Keywords: | Audiovisual active speaker detection; audiovisual diarisation; speech processing; egocentric recordings; augmented reality; |
| Awarding institution: | University of Sheffield |
| Academic Units: | The University of Sheffield > Faculty of Engineering (Sheffield) > Computer Science (Sheffield) |
| Date Deposited: | 15 Jun 2026 09:56 |
| Last Modified: | 15 Jun 2026 09:56 |
| Open Archives Initiative ID (OAI ID): | oai:etheses.whiterose.ac.uk:38775 |
Download
Final eThesis - complete (pdf)
Filename: Jason_Clarke_PhD_thesis.pdf
Licence:

This work is licensed under a Creative Commons Attribution NonCommercial NoDerivatives 4.0 International License
Export
Statistics
You do not need to contact us to get a copy of this thesis. Please use the 'Download' link(s) above to get a copy.
You can contact us about this thesis. If you need to make a general enquiry, please see the Contact us page.