Panahi, Solmaz (2026) Bridging the Data Gap: LLM-Driven Synthetic Data Generation for Low-Resource Natural Language Inference. PhD thesis, National University of Ireland Maynooth.
Preview
Available under License Creative Commons Attribution Non-commercial Share Alike.
Download (7MB) | Preview
Abstract
This thesis addresses the challenge of data scarcity in low-resource languages by leveraging
the dual role of Large Language Models (LLMs) to both generate and annotate data.
We evaluate the reliability of LLM-based annotations and examine the effectiveness of
LLM-generated synthetic data in improving downstream task performance. This research
uses the Natural Language Inference (NLI) task in Farsi as its case study. Our approach
involves benchmarking both open-weight and closed-source LLMs on a gold-standard
dataset, developing an LLM-driven pipeline to generate labeled premise-hypothesis pairs,
and training mT5 models to measure performance gains. Further experiments were conducted
with various LLM annotators to analyze how annotation quality affects model
outcomes.
The results reveal key insights. First, synthetic data substantially improves model
performance, particularly under sequential training regimes, confirming its effectiveness
as a strategy for data augmentation in low-resource settings. Second, the quality of annotations
plays a decisive role in downstream outcomes—higher-quality labels consistently
lead to greater accuracy and robustness. Finally, the reliability of an LLM as an annotator
emerges as a multifaceted property, shaped by factors including model architecture,
prompt design, and the inherent complexity of the task itself.
These findings show that while LLM-generated synthetic data can mitigate data
scarcity issues, annotation quality remains a critical factor influencing model performance.
The thesis contributes insights into optimizing annotation pipelines and synthetic data
utility for low-resource NLP, with implications for multilingual model training strategies.
| Item Type: | Thesis (PhD) |
|---|---|
| Keywords: | Bridging; data gap; LLM-Driven Synthetic Data Generation; Low-Resource; Natural Language Inference; |
| Academic Unit: | Faculty of Science & Engineering > Research Institutes > Hamilton Institute |
| Item ID: | 21865 |
| Depositing User: | IR eTheses |
| Date Deposited: | 28 Sep 2026 12:24 |
| Use Licence: | This item is available under a Creative Commons Attribution Non Commercial Share Alike Licence (CC BY-NC-SA). Details of this licence are available here |
Downloads
Downloads per month over past year
Share and Export
Share and Export