TOLID: Turkish Offensive Language Identification Dataset and Transformer-Based Benchmarks
Electrica, cilt.26, 2026 (ESCI, Scopus, TRDizin)
- Yayın Türü: Makale / Tam Makale
- Cilt numarası: 26
- Basım Tarihi: 2026
- Doi Numarası: 10.5152/electrica.2026.25404
- Dergi Adı: Electrica
- Derginin Tarandığı İndeksler: Emerging Sources Citation Index (ESCI), Scopus, TR DİZİN (ULAKBİM)
- Anahtar Kelimeler: Deep learning, hate speech, offensive language, text classification, Turkish
- Açık Arşiv Koleksiyonu: AVESİS Açık Erişim Koleksiyonu
- İstanbul Üniversitesi-Cerrahpaşa Adresli: Evet
Özet
This study presents the Turkish Offensive Language Identification Dataset (TOLID), a large-scale, high-quality dataset for the automatic detection of offensive language in Turkish social media posts and evaluates the performance of several transformer-based models. Although most previous studies have been constrained by small sample sizes, imbalanced class distributions, or narrow topical focus, TOLID includes a wide range of offensive expressions without imposing restrictions on topics, individuals, or groups. The dataset was annotated by three independent experts using a hierarchical and fine-grained scheme that addresses not only general offensive language but also specific subcategories such as sexist, racist, political, and religious insults. To evaluate the dataset, several transformer-based models adapted for Turkish, including BERTurk (Bidirectional Encoder Representations from Transformers for Turkish), ConvBERTurk (Convolutional Bidirectional Encoder Representations from Transformers for Turkish), and ELECTRA-Turkish (Efficiently Learning an Encoder that Classifies Token Replacements Accurately for Turkish), were trained and tested. Among these, ConvBERTurk achieved the highest scores, reaching 82.76% macro F1 in offensive vs. non-offensive classification and 78.14% in targeted vs. non-targeted offensive classification, outperforming previous research. These results demonstrate that combining a balanced, multi-annotated dataset with advanced deep learning architectures can effectively address the linguistic richness, contextual complexity, and informal nature of Turkish social media text. Additionally, a web-based application was developed to provide a practical interface for analyzing text and visualizing model outputs, extending the study's impact beyond academic research to real-world applications. Overall, this study addresses key limitations of prior research and makes a significant contribution to Turkish natural language processing by providing a meticulously constructed dataset and extensive benchmarks with state-of-the-art models. TOLID establishes a robust foundation for future work on offensive language subtypes, automatic moderation, and toxicity analysis in Turkish social media.