Comparative Analysis of CNN-LSTM and LSTM Models for Cyberbullying Detection with Increasing Dataset Sizes

 Tri Pratiwi Handayani, Mohamad Ilyas Abas

Abstract


This study compares the performance of two deep learning models, CNN-LSTM and LSTM, for identifying cyberbullying in social media text. Three distinct dataset sizes are used for our evaluation and comparison: 1,000, 5,000, and 10,000 samples. The results indicate that the CNN-LSTM model outperforms the LSTM-only model (Ablation model) for the largest dataset size, exhibiting substantial enhancements in accuracy, precision, recall, and F1-Score as the dataset size increases. The Ablation model exhibits competitive performance and slightly superior results on the mid-sized dataset. However, it inevitably falls behind the CNN-LSTM model when trained on 10,000 samples. These findings imply that increasing the complexity of the CNN layer in the CNN-LSTM model improves its ability to collect significant features in bigger datasets, making it more successful for cyberbullying detection.

 


Full Text:

PDF

References


C. Van Hee et al., “Automatic detection of cyberbullying in social media text,” PLoS One, vol. 13, no. 10, 2018, doi: 10.1371/journal.pone.0203794.

H. Dani, J. Li, and H. Liu, “Sentiment Informed Cyberbullying Detection in Social Media,” in Lecture Notes in Computer Science (including subseries Lecture Notes in Artificial Intelligence and Lecture Notes in Bioinformatics), 2017. doi: 10.1007/978-3-319-71249-9_4.

M. A. Darmawan, N. William Boentoro, K. C. Surya, and R. Sutoyo, “Experiments on IndoBERT Implementation for Detecting Multi-Label Hate Speech with Data Resampling through Synonym Replacement Method,” in 8th International Conference on Recent Advances and Innovations in Engineering: Empowering Computing, Analytics, and Engineering Through Digital Innovation, ICRAIE 2023, 2023. doi: 10.1109/ICRAIE59459.2023.10468099.

M. Kamyab, G. Liu, and M. Adjeisah, “Attention-Based CNN and Bi-LSTM Model Based on TF-IDF and GloVe Word Embedding for Sentiment Analysis,” Appl. Sci., vol. 11, no. 23, 2021, doi: 10.3390/app112311255.

H. Imaduddin, L. A. Kusumaningtias, and F. Y. A’la, “Application of LSTM and GloVe Word Embedding for Hate Speech Detection in Indonesian Twitter Data,” Ing. des Syst. d’Information, vol. 28, no. 4, pp. 1107–1112, 2023, doi: 10.18280/isi.280430.

L. Anindyati, A. Purwarianti, and A. Nursanti, “Optimizing Deep Learning for Detection Cyberbullying Text in Indonesian Language,” in Proceedings - 2019 International Conference on Advanced Informatics: Concepts, Theory, and Applications, ICAICTA 2019, 2019. doi: 10.1109/ICAICTA.2019.8904108.

T. P. Handayani and H. Gani, “Enhancing Multi-Label Hate Speech and Abusive Language Detection on Indonesian Twitter Using Recurrent Neural Networks with Hyperparameter Tuning,” J. Ilm. Tek. Mesin, Elektro dan Komput., vol. 3, no. 3, pp. 602–612, Nov. 2023, [Online]. Available: https://scholar.google.com/scholar?cluster=17876321117069128543

T. P. Handayani, W. Hasyim, and Nursetia, “Preliminary Evaluation of Gaussian Naive Bayes for Multi-Label Hate Speech and Abusive Language Detection on Indonesian Twitter,” J. Int. Multidiscip. Res., vol. 1, no. 1, pp. 1–7, Nov. 2023, [Online]. Available: https://scholar.google.com/scholar?cluster=17876321117069128544

M. A. A. Yani and W. Maharani, “Analyzing Cyberbullying Negative Content on Twitter Social Media with the RoBERTa Method,” JINAV J. Inf. …, 2023, [Online]. Available: https://jpabdimas.idjournal.eu/index.php/jinav/article/view/1543

A. Muhariya, I. Riadi, Y. Prayudi, and ..., “Utilizing K-means Clustering for the Detection of Cyberbullying Within Instagram Comments.,” … des Systèmes d’ …, 2023, [Online]. Available: https://search.ebscohost.com/login.aspx?direct=true&profile=ehost&scope=site&authtype=crawler&jrnl=16331311&AN=171938951&h=jAR8LFD6hrLxqT5pNnTroYYiMkgwU7GV9uo%2FcpMzrDkpFvWMbDFoaiuRfFKcJG1xL%2BHRfUycy2tq0OOGAZ6WEA%3D%3D&crl=c

N. I. Purnayasa, I. M. A. D. Suarjaya, and I. P. A. Dharmaadi, “Analysis of Cyberbullying Level using Support Vector Machine Method,” J. Ilm. Merpati (Menara Penelit. Akad. Teknol. Informasi), vol. 10, no. 2, 2022, doi: 10.24843/jim.2022.v10.i02.p01.

J. Y. Anugrah and N. M. Aesthetika, “Teenagers’ Perception of Cyberbullying on Instagram,” KnE Soc. Sci., 2022, doi: 10.18502/kss.v7i12.11532.

T. Nugraha Manoppo and D. Hatta Fudholi, “Deteksi Cyberbullying berdasarkan Unsur Perbuatan Pidana yang Dilanggar dengan Naive Bayes dan Support Vector Machine,” J. Sains Komput. Inform. (J-SAKTI, vol. 5, no. 1, 2021.

A. Syahid, D. Sudana, and A. D. Bachari, “Cyberbullying on Social Media in Indonesia and Its Legal Impact: Analysis of Language Use in Ethnicity, Religious, Racial, and Primordial Issues,” Theory Pract. Lang. Stud., vol. 13, no. 8, 2023, doi: 10.17507/tpls.1308.09.

M. O. Ibrohim and I. Budi, “Multi-label Hate Speech and Abusive Language Detection in Indonesian Twitter,” ALW3 3rd Work. Abus. Lang. Online, pp. 46–57, 2019, [Online]. Available: https://www.aclweb.org/anthology/W19-3506.pdf


DOI: http://dx.doi.org/10.31314/juik.v4i2.3185

Article metrics

Abstract views : 352 | views : 231

Refbacks

  • There are currently no refbacks.


Copyright (c) 2024 Jurnal Ilmu Komputer (JUIK)

Creative Commons License
This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.