Publication Date
2026
Document Type
Dissertation/Thesis
First Advisor
Alhoori, Hamed
Degree Name
M.S. (Master of Science)
Legacy Department
Department of Computer Science
Abstract
The landscape of scientific communication has undergone a significant transformation with the emergence and widespread adoption of Large Language Models (LLMs). LLMs have enhanced productivity and creativity in scientific writing, raising serious questions about the faithfulness and reliability of the generated content. A common problem in LLMs is hallucination, which occurs when models generate fluent but inaccurate or inconsistent content. When hallucinations enter scientific writing, the intended meaning may be distorted, coherence disrupted, and the study’s overall integrity threatened.
This thesis investigates hallucinations in scientific writing from both an analytical and dataset-driven approach. First, it offers a comprehensive analysis of contextual inconsistency, a type of faithfulness hallucination in research papers on Natural Language Processing (NLP) from two different eras: before and after LLMs became widely adopted. Using NLP as a representative domain, the study looks at how context-related contradictions change over time, assesses how serious they are at the paragraph level, and considers how they might be related to writing that is aided by LLM. The results show that contextual inconsistencies have become more common in recent years, and it is getting harder to distinguish AI-generated text from human-written content.
Guided by the findings of the trend analysis from our first study, the thesis introduces SciHallu, a multi-granularity dataset designed for hallucination identification in scientific writing. SciHallu covers multiple academic fields, such as Computer Science, Health Sciences, and Humanities and Social Sciences. The dataset captures the layered nature of hallucinations, facilitating fine-grained examination of hallucination at the token, sentence, and paragraph levels. The instances are built by applying controlled perturbations to source texts from pre-LLM research articles. For training and assessment purposes, each instance is accompanied by a rationale explaining the perturbation details. Expert human annotation is used to validate the dataset.
Finally, the results obtained using the SciHallu benchmark demonstrate that existing models struggle in correctly identifying hallucinations in scientific texts, particularly when they are subtle or occur at lower granular levels. These findings highlight the need for specialized resources to address the unique challenges related to scientific hallucinations and indicate a crucial gap in the current LLM capabilities.
Collectively, this thesis offers a thorough analysis of hallucinatory behavior present in scientific writing and serves as a benchmark resource to aid in the creation of more dependable, and trustworthy language models for academic communication.
Recommended Citation
Hossain, Adiba Ibnat, "Hallucination Detection in Scientific Writing: Quantification, Trend Analysis and Benchmarking" (2026). Graduate Research Theses & Dissertations. 8211.
https://huskiecommons.lib.niu.edu/allgraduate-thesesdissertations/8211
Extent
100 pages
Language
en
Publisher
Northern Illinois University
Rights Statement
In Copyright
Rights Statement 2
NIU theses are protected by copyright. They may be viewed from Huskie Commons for any purpose, but reproduction or distribution in any format is prohibited without the written permission of the authors.
Media Type
Text
