Year

2026

Season

Summer

Paper Type

Master's Thesis

College

College of Computing, Engineering & Construction

Degree Name

Master of Science in Computer and Information Sciences (MS)

Department

Computing

Committee Chairperson

Dr. Indika Kahanda

Second Advisor

Dr. Andrea Arikawa

Third Advisor

Dr. Sandeep Reddivari

Department Chair

Dr.Nan Niu

College Dean

Dr.William Klostermeyer

Abstract

Nutrition misinformation on social media often arises from selective or decontextualized interpretations of established dietary evidence, particularly in contested domains such as seed oils and omega-6 fatty acids. To address this challenge, we present a unified framework for annotation, detection, and evidence-grounded interpretation of nutrition-related content on Instagram, integrating large language models (LLMs), interpretable machine learning, and retrieval-augmented generation (RAG) grounded in computable biomedical knowledge. We curate a dataset of 317 Instagram captions, consisting of 169 seed oil-related and 148 omega-6-related posts. Expert nutritionists independently annotate posts using the U.S. Dietary Guidelines (2020–2025), producing gold-standard labels via majority voting. We evaluate five open-source LLMs as automated annotators and introduce a hierarchical error taxonomy that categorizes misclassifications by direction, mechanism, and contributing factors, revealing systematic failure modes such as misinterpretation of nuanced claims and overconfident reasoning. In parallel, we benchmark three complementary misinformation detection approaches: feature-based traditional machine learning models that incorporate linguistic, rhetorical, affective, and psychological signals; embedding-based classification; and transformer-based fine tuning. We further assess model robustness under in-domain and cross-domain set tings, highlighting differences in generalization between semantic and linguistically grounded representations. Transformer-based models achieve the strongest performance in in-domain settings, while feature-based representations remain competitive and offer improved interpretability under constrained evaluation conditions. Finally, we develop a RAG framework that operationalizes the U.S. Dietary Guide lines as a sentence-indexed knowledge base. Retrieved evidence is paired with each post to enable a language model to generate grounded classifications and explanations that explicitly align or contrast user claims with authoritative dietary recommendations. Results indicate that retrieval grounding in the U.S. Dietary Guidelines improves factual consistency and the quality of model-generated explanations. This work has implications for the design of reliable and interpretable health misinformation detection systems that integrate human expertise, machine learning models, and structured biomedical knowledge.

Available for download on Thursday, July 15, 2027

Share

COinS
 

Accessibility Statement

This item was created or digitized before April 24, 2027, or is a reproduction of legacy material created before that date. It is preserved in its original, unmodified state specifically for research, reference, or historical recordkeeping. In accordance with the ADA Title II Final Rule, the Library provides accessible versions of archival materials by request. If you are experiencing difficulty accessing the information on the site due to a disability, please submit a request through the following form for assistance.