Turkish and English disinformation detection using deep learning and large language models: Introducing the Dogrulamac dataset


Ayaydin O. B., Orhan T., Orakci A. A., Sagdic S., GÜVEN E. Y.

EXPERT SYSTEMS WITH APPLICATIONS, cilt.333, 2027 (SCI-Expanded, Scopus)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 333
  • Basım Tarihi: 2027
  • Doi Numarası: 10.1016/j.eswa.2026.133847
  • Dergi Adı: EXPERT SYSTEMS WITH APPLICATIONS
  • Derginin Tarandığı İndeksler: Science Citation Index Expanded (SCI-EXPANDED), Scopus, Aerospace Database, Applied Science & Technology Source, Compendex, INSPEC, Public Affairs Index, Academic Search Ultimate (EBSCO), Engineering Source (EBSCO), Technology Collection (ProQuest)
  • İstanbul Üniversitesi-Cerrahpaşa Adresli: Evet

Özet

The proliferation of disinformation across online platforms poses a significant threat to information integrity, particularly in multilingual contexts. Although manual fact-checking seeks to identify false information, the rapid flow of social media content exceeds the capacity of human verification efforts. This study presents a comprehensive approach to automated disinformation detection using both English and Turkish resources. We introduce Dogrulamac, a novel Turkish disinformation dataset designed to address the scarcity of labeled data in low-resource languages and to support reproducible multilingual NLP research. For the English language, a comprehensive disinformation dataset was rigorously deduplicated to prevent data leakage, yielding a highly balanced corpus of 68,605 unique articles. To ensure statistical robustness and mitigate partition variance, all deep learning architectures including LSTM, BERT, and the large language model Gemma-2 were evaluated using a 5-fold cross-validation strategy. The models were compared against each other and with prior findings reported in the literature. Experimental results demonstrate that Gemma-2 achieved the highest average accuracy, with 97.39% +/- 0.15% on the Turkish dataset and 98.15% +/- 0.10% on the English dataset. These rigorously validated findings confirm the effectiveness of transformer-based architectures, particularly Gemma-2, in cross-lingual disinformation detection and establish strong, statistically stable baselines for future multilingual NLP research.