TOLID: Turkish Offensive Language Identification Dataset and Transformer-Based Benchmarks


Creative Commons License

KURT M. S., YÜCEL DEMİREL E.

Electrica, cilt.26, 2026 (ESCI, Scopus, TRDizin)

  • Yayın Türü: Makale / Tam Makale
  • Cilt numarası: 26
  • Basım Tarihi: 2026
  • Doi Numarası: 10.5152/electrica.2026.25404
  • Dergi Adı: Electrica
  • Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, TR DİZİN (ULAKBİM)
  • Anahtar Kelimeler: Deep learning, hate speech, offensive language, text classification, Turkish
  • Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
  • İstanbul Üniversitesi-Cerrahpaşa Adresli: Evet

Özet

This study presents the Turkish Offensive Language Identification Dataset (TOLID), a large-scale, high-quality dataset for the automatic detection of offensive language in Turkish social media posts and evaluates the performance of several transformer-based models. Although most previous studies have been constrained by small sample sizes, imbalanced class distributions, or narrow topical focus, TOLID includes a wide range of offensive expressions without imposing restrictions on topics, individuals, or groups. The dataset was annotated by three independent experts using a hierarchical and fine-grained scheme that addresses not only general offensive language but also specific subcategories such as sexist, racist, political, and religious insults. To evaluate the dataset, several transformer-based models adapted for Turkish, including BERTurk (Bidirectional Encoder Representations from Transformers for Turkish), ConvBERTurk (Convolutional Bidirectional Encoder Representations from Transformers for Turkish), and ELECTRA-Turkish (Efficiently Learning an Encoder that Classifies Token Replacements Accurately for Turkish), were trained and tested. Among these, ConvBERTurk achieved the highest scores, reaching 82.76% macro F1 in offensive vs. non-offensive classification and 78.14% in targeted vs. non-targeted offensive classification, outperforming previous research. These results demonstrate that combining a balanced, multi-annotated dataset with advanced deep learning architectures can effectively address the linguistic richness, contextual complexity, and informal nature of Turkish social media text. Additionally, a web-based application was developed to provide a practical interface for analyzing text and visualizing model outputs, extending the study's impact beyond academic research to real-world applications. Overall, this study addresses key limitations of prior research and makes a significant contribution to Turkish natural language processing by providing a meticulously constructed dataset and extensive benchmarks with state-of-the-art models. TOLID establishes a robust foundation for future work on offensive language subtypes, automatic moderation, and toxicity analysis in Turkish social media.