Alahmari, Salwa Saad M (2026) A Corpus-Based Approach to Saudi Arabian Dialects: Implications for Dialect Identification, Machine Translation and Sentiment Analysis. PhD thesis, University of Leeds.
Abstract
Arabic is a linguistically diverse language comprising numerous dialects spoken across
the Arab world, with the dialects of Saudi Arabia representing a particularly rich
and complex landscape shaped by regional, historical, and social influences. Despite
their linguistic and cultural importance, Saudi Arabic dialects remain significantly
underrepresented in computational linguistics and natural language processing (NLP)
research and are often treated as low-resource language varieties due to the scarcity of
annotated data, standardized resources, and evaluation benchmarks. Existing resources
and models predominantly focus on Modern Standard Arabic (MSA) or broad regional
dialect groups, such as Egyptian or Gulf Arabic, leaving fine-grained Saudi dialectal
variation largely unexplored. This gap limits the development of dialect-aware NLP
systems capable of accurately processing Saudi Arabian dialects.
The main aim of this thesis is to systematically investigate fine-grained Saudi Arabic
dialects across three core NLP tasks, dialect identification, machine translation, and
sentiment analysis, by evaluating the performance of both pre-trained language models
(PLMs) and large language models (LLMs) in processing low-resource Saudi dialectal
data. To achieve this aim, the study introduces several novel datasets designed to
support the analysis of Saudi Arabian dialects across different NLP tasks. For dialect
identification, the Saudi Arabian Tweets Corpus (SATC) is developed, comprising Saudi
Arabic tweets representing five major regional dialects. In addition, the Saudi Arabian
Dialectal Song Lyrics Corpus (SADSLyC) is constructed based on Saudi song lyrics,
capturing culturally rich and regionally diverse linguistic expressions. For machine
translation, the study presents SADSLyC-E-MSA, a parallel subset of the SADSLyC
corpus that includes aligned translations in English and Modern Standard Arabic.
Finally, for sentiment analysis, the research introduces the Saudi Arabian Proverbs Corpus (SAPC), a collection of authentic proverbs gathered from different regions of
Saudi Arabia, reflecting deeply rooted cultural meanings and regional dialectal variation.
The findings demonstrate that task-specific fine-tuning combined with dialect-
focused corpora significantly improves language models performance on fine-grained
Saudi dialect in different NLP tasks. In the machine translation experiments, the best-
performing model achieved an improvement of 39.17% over the baseline model without
fine-tuning when translating Saudi Arabian dialects, highlighting the substantial impact
of adapting models to dialect-specific data.
Metadata
| Supervisors: | Atwell, Eric and Alsalka, Ammar and Saadany, Hadeel |
|---|---|
| Related URLs: | |
| Keywords: | Arabic NLP; Dialectal Arabic; Saudi Dialects; Dialect identification; Machine Translation; Sentiment Analysis |
| Awarding institution: | University of Leeds |
| Academic Units: | The University of Leeds > Faculty of Engineering (Leeds) > School of Computing (Leeds) |
| Date Deposited: | 23 Jul 2026 14:24 |
| Last Modified: | 23 Jul 2026 14:24 |
| Open Archives Initiative ID (OAI ID): | oai:etheses.whiterose.ac.uk:39095 |
Download
Final eThesis - complete (pdf)
Embargoed until: 1 August 2031
Please use the button below to request a copy.
Filename: ALAHMARI_SSM_ComputerScience_PhD_2026.pdf
Export
Statistics
Please use the 'Request a copy' link(s) in the 'Downloads' section above to request this thesis. This will be sent directly to someone who may authorise access.
You can contact us about this thesis. If you need to make a general enquiry, please see the Contact us page.