MURAL - Maynooth University Research Archive Library



    Bridging the Data Gap: LLM-Driven Synthetic Data Generation for Low-Resource Natural Language Inference


    Panahi, Solmaz (2026) Bridging the Data Gap: LLM-Driven Synthetic Data Generation for Low-Resource Natural Language Inference. PhD thesis, National University of Ireland Maynooth.

    Abstract

    This thesis addresses the challenge of data scarcity in low-resource languages by leveraging the dual role of Large Language Models (LLMs) to both generate and annotate data. We evaluate the reliability of LLM-based annotations and examine the effectiveness of LLM-generated synthetic data in improving downstream task performance. This research uses the Natural Language Inference (NLI) task in Farsi as its case study. Our approach involves benchmarking both open-weight and closed-source LLMs on a gold-standard dataset, developing an LLM-driven pipeline to generate labeled premise-hypothesis pairs, and training mT5 models to measure performance gains. Further experiments were conducted with various LLM annotators to analyze how annotation quality affects model outcomes. The results reveal key insights. First, synthetic data substantially improves model performance, particularly under sequential training regimes, confirming its effectiveness as a strategy for data augmentation in low-resource settings. Second, the quality of annotations plays a decisive role in downstream outcomes—higher-quality labels consistently lead to greater accuracy and robustness. Finally, the reliability of an LLM as an annotator emerges as a multifaceted property, shaped by factors including model architecture, prompt design, and the inherent complexity of the task itself. These findings show that while LLM-generated synthetic data can mitigate data scarcity issues, annotation quality remains a critical factor influencing model performance. The thesis contributes insights into optimizing annotation pipelines and synthetic data utility for low-resource NLP, with implications for multilingual model training strategies.
    Item Type: Thesis (PhD)
    Keywords: Bridging; data gap; LLM-Driven Synthetic Data Generation; Low-Resource; Natural Language Inference;
    Academic Unit: Faculty of Science & Engineering > Research Institutes > Hamilton Institute
    Item ID: 21865
    Depositing User: IR eTheses
    Date Deposited: 28 Sep 2026 12:24
    Use Licence: This item is available under a Creative Commons Attribution Non Commercial Share Alike Licence (CC BY-NC-SA). Details of this licence are available here

    Downloads

    Downloads per month over past year

    Origin of downloads

    Repository Staff Only (login required)

    Item control page
    Item control page