Meghanani, Amit
ORCID: 0000-0002-0811-274X
(2026)
Improving Content Representations in Self-supervised Speech Models.
PhD thesis, University of Sheffield.
Abstract
Despite progress in language technologies, with large language models achieving remarkable results on text, modelling speech remains difficult. Speech carries linguistic content together with many other signals, including speaker identity, emotion, and background noise, making it harder to model than discrete text. Although prior work has explored disentangling these factors, strict unsupervised separation remains unresolved. Meanwhile, Self-supervised Learning (SSL) has transformed speech processing by learning rich representations from unlabelled data. However, these representations are not fully task-agnostic, and improving them for particular downstream tasks can require task-aware pre-training or substantial labelled data.
This thesis investigates how content-related information in SSL-based speech representations can be strengthened cost-effectively while reducing the influence of less relevant factors. Rather than aiming for complete disentanglement, it studies Self-supervised Fine-tuning (SSFT) as an intermediate stage between SSL pre-training and downstream fine-tuning to make representations more suitable for content-focused tasks. It first shows that correspondence training, originally developed for Mel-frequency Cepstral Coefficient (MFCC)-based Acoustic Word Embeddings (AWEs), can be successfully applied to SSL-based representations, suggesting that their content-related information can be improved through lightweight adaptation rather than expensive pre-training from scratch.
The thesis then introduces Contrastive Learning with Only Positive Pairs for Speech (CLOPS) as a unifying framework for SSFT. Within this framework, correspondence-style learning is extended from fixed word segments to full utterance-level sequences, with two CLOPS-based methods, based on fixed reference regularisation and temporal regularisation, shown to improve content-related downstream tasks efficiently. The framework is further extended to noisy conditions to obtain more noise-robust representations, while analysis of frame-wise and alignment-aware objectives highlights position-related effects. Experiments on SUPERB show that CLOPS-based methods improve content relevance and robustness at a fraction of the cost of task-aware pre-training from scratch, offering a unified and practical perspective on deriving content-focused speech representations from general-purpose SSL models.
Metadata
| Supervisors: | Hain, Thomas |
|---|---|
| Awarding institution: | University of Sheffield |
| Academic Units: | The University of Sheffield > Faculty of Engineering (Sheffield) > Computer Science (Sheffield) |
| Date Deposited: | 15 Jun 2026 09:58 |
| Last Modified: | 15 Jun 2026 09:58 |
| Open Archives Initiative ID (OAI ID): | oai:etheses.whiterose.ac.uk:38761 |
Download
Final eThesis - complete (pdf)
Embargoed until: 13 May 2027
Please use the button below to request a copy.
Filename: Amit_Meghanani_Thesis.pdf
Export
Statistics
Please use the 'Request a copy' link(s) in the 'Downloads' section above to request this thesis. This will be sent directly to someone who may authorise access.
You can contact us about this thesis. If you need to make a general enquiry, please see the Contact us page.