A Novel Integration of Exploratory Data Analysis and LDA Text Mining for Deeper Understanding of Libyan-Related Tweets
Abstract
Identifying topics across digital sources such as scientific articles, web publications, and social media platforms is essential for summarising large text datasets and understanding emerging trends. Social networking platforms, in particular, generate vast amounts of data, creating an urgent need for topic extraction to uncover underlying themes. This study aims to extract topics from Arabic-language tweets posted by Libyan users on Twitter using a structured multi-stage approach. The study follows a four-stage methodology: (1) Data Collection, involving the construction of a Twitter corpus based on Libyan dialect-specific keywords; (2) Exploratory Analysis, which employs statistical analysis to identify key patterns in the data; (3) Preprocessing, which applies natural language processing techniques to prepare the dataset; and (4) Topic Modelling, where Latent Dirichlet Allocation (LDA) is applied to extract dominant topics. The exploratory analysis revealed behavioural patterns such as a high frequency of positive emoji usage, peak Twitter activity at 9 p.m., and the dominance of Android as the operating system. The topic modelling results identified major themes including Friendship, Travel, Political Issues, and Education. These findings provide insights into the interests and communication patterns of Libyan Twitter users and demonstrate the effectiveness of topic modelling in capturing culturally and regionally relevant discourse in social media data.
Full text article
References
[1] Q. Shen, ‘Topic Discovery and Future Trend Prediction In Scholarly Networks’. 2016.
[2] T. Porturas and R. A. Taylor, ‘Forty years of emergency medicine research: Uncovering research themes and trends through topic modeling’, Am. J. Emerg. Med., vol. 45, pp. 213–220, 2021.
[3] R. Abdelhamed, M. Essgaer, A. Agaal, and A. Shibani, ‘Enhancing Topic Modeling in Scientific Literature: A Comparative Study of LDA with Word2Vec, Doc2Vec, and SciBERT Embeddings’, in Selected Papers from the International Conference on Artificial Intelligence, A. O. Albaji, Ed., Cham: Springer Nature Switzerland, 2026, pp. 634–648.
[4] R. Abdelhamed and M. Essgaer, ‘Exploring Knowledge Landscapes: Clustering and Topic Modeling of Sebha University Scientific Publications’, in 2024 IEEE 4th International Maghreb Meeting of the Conference on Sciences and Techniques of Automatic Control and Computer Engineering (MI-STA), May 2024, pp. 712–717. doi: 10.1109/MI-STA61267.2024.10599683.
[5] M. B. Mutanga and A. Abayomi, ‘Tweeting on COVID-19 pandemic in South Africa: LDA-based topic modelling approach’, Afr. J. Sci. Technol. Innov. Dev., vol. 14, no. 1, pp. 163–172, 2022.
[6] Y. Chen and L. Liu, ‘Development and research of topic detection and tracking’, in 2016 7th IEEE International Conference on Software Engineering and Service Science (ICSESS), IEEE, 2016.
[7] A. Krishnan, ‘Exploring the Power of Topic Modeling Techniques in Analyzing Customer Reviews: A Comparative Analysis’. 2023.
[8] et al Albalawi R., ‘Using topic modeling methods for short-text data: A comparative analysis’, Front. Artif. Intell., vol. 3, p. 42, 2020.
[9] R. Egger, ‘Topic Modelling’, Appl. Data Sci. Tour., 2022.
[10] S. Kumar and V. Bhatnagar, ‘A review of regression models in machine learning’, J. Intell. Syst. Comput., vol. 3, no. 1, pp. 40–47, 2022.
[11] et al Bohra N., ‘Popularity Prediction of Social Media Post Using Tensor Factorization’, Intell. Autom. Soft Comput., vol. 36, no. 1, 2023.
[12] A. andL. B. R. Meddeb, ‘Using Topic Modeling and Word Embedding for Topic Extraction in Twitter’, in International Conference on Knowledge-Based Intelligent Information & Engineering Systems, 2022.
[13] et al Alothman M., ‘Review of Researches on Arabic Social Media Text Mining’. 2021.
[14] N. F. Bin Hathlian and A. M. Hafezs, ‘Sentiment - subjective analysis framework for arabic social media posts’, in 2016 4th Saudi International Conference on Information Technology (Big Data Analysis) (KACSTIT): 1-6, 2016.
[15] N. F. B. Hathlian and A. M. Hafez, ‘Subjective Text Mining for Arabic Social Media’, Int J Semantic Web Inf Syst, vol. 13, pp. 1–13, 2017.
[16] L. Al-Horaibi and M. B. Khan, ‘Sentiment analysis of Arabic tweets using text mining techniques’, in International Workshop on Pattern Recognition, 2016.
[17] et al Ghani N. A., ‘Social media big data analytics: A survey’, Comput. Hum. Behav., vol. 101, pp. 417–428, 2019.
[18] et al Alqmase M., ‘Sports-fanaticism formalism for sentiment analysis in Arabic text’, Soc. Netw. Anal. Min., vol. 11, no. 1, pp. 1–24, 2021.
[19] et al Asif M., ‘Sentiment analysis of extremism in social media from textual information’, Telemat. Inform., vol. 48, p. 101345, 2020.
[20] S. C. McGregor, ‘Social media as public opinion: How journalists use social media to represent public opinion’, Journalism, vol. 20, no. 8, pp. 1070–1086, 2019.
[21] et al Hung M., ‘Social network analysis of COVID-19 sentiments: Application of artificial intelligence’, J. Med. Internet Res., vol. 22, no. 8, 2020.
[22] et al Ragini J. R., ‘Big data analytics for disaster response and recovery through sentiment analysis’, Int. J. Inf. Manag., vol. 42, pp. 13–24, 2018.
[23] S. A. AlAjlan and A. K. J. Saudagar, ‘Machine learning approach for threat detection on social media posts containing Arabic text’, Evol. Intell., vol. 14, no. 2, pp. 811–822, 2021.
[24] et al Khalafat M., ‘Violence Detection over Online Social Networks: An Arabic Sentiment Analysis Approach’, iJIM, vol. 15, no. 14, p. 91, 2021.
[25] et al Kanan T., ‘Cyber-bullying and cyber-harassment detection using supervised machine learning techniques in Arabic social media contents’, J. Internet Technol., vol. 21, no. 5, pp. 1409–1421, 2020.
[26] et al Guellil I., ‘Detecting hate speech against politicians in Arabic community on social media’, Int. J. Web Inf. Syst., 2020.
[27] etal Allaith A., ‘Neural Network Approach for Irony Detection from Arabic Text on Social Media’, in FIRE (Working Notes), 2019.
[28] T. Al-Khalifi, ‘Social media data mining and its applications in media research: Sentiment analysis as a model’, J. Media Res. Stud., vol. 8, no. 8, pp. 1–73, 2019.
[29] et al Curiskis S. A., ‘An evaluation of document clustering and topic modelling in two online social networks: Twitter and Reddit’, Inf. Process. Manag., vol. 57, no. 2, p. 102034, 2020.
[30] E. F. E. Elnour, ‘Using Data Mining Techniques to Establish Standard Sizing System for Sudanese Army Officers Poshirt’. 2018.
[31] et al Yuan M., ‘Document Clustering vs Topic Models: A Case Study’, in Proceedings of the 25th Australasian Document Computing Symposium, 2021.
[32] et al Lossio-Ventura J. A., ‘Evaluation of clustering and topic modeling methods over health-related tweets and emails’, Artif. Intell. Med., vol. 117, p. 102096, 2021.
[33] M. H. Beseiso, ‘New Sentiment Analysis Model Using LDA for Arabic Tweets’, in Proceedings of the 3rd International Conference on Advances in Artificial Intelligence, 2019.
[34] et al Al-Qudah I., ‘Applying Latent Dirichlet Allocation Technique to Classify Topics on Sustainability Using Arabic Text’, Sai, 2022.
[35] et al Albadarneh J., ‘Using Big Data Analytics for Authorship Authentication of Arabic Tweets’, in 2015 IEEE/ACM 8th International Conference on Utility and Cloud Computing (UCC): 448-452, 2015.
[36] et al Zrigui M., ‘Arabic Text Classification Framework Based on Latent Dirichlet Allocation’, J Comput Inf Technol, vol. 20, pp. 125–140, 2012.
[37] et al Wang H., ‘Exploring the Chinese public’s perception of omicron variantson social media: lda-based topic modeling and sentiment analysis’, Int. J. Environ. Res. Public. Health, vol. 19, no. 14, p. 8377, 2022.
[38] L. Zou and W. W. Song, ‘LDA-TM: A two-step approach to Twitter topic data clustering’, in 2016 IEEE International Conference on Cloud Computing and Big Data Analysis (ICCCBDA): 342-347, 2016.
[39] S. Yang and H. Zhang, ‘Text Mining of Twitter Data Using a Latent Dirichlet Allocation Topic Modeland Sentiment Analysis’. 2018.
[40] A. Omar, M. Essgaer, and K. M. S. Ahmed, ‘Using Machine Learning Model To Predict Libyan Telecom Company Customer Satisfaction’, in 2022 International Conference on Engineering & MIS (ICEMIS), Jul. 2022, pp. 1–6. doi: 10.1109/ICEMIS56295.2022.9914055.
[41] et al Oshikawa R., ‘A survey on natural language processing for fake news detection’. 2018.
[42] et al Hegazi M. O., ‘Preprocessing Arabic text on social media’, Heliyon, vol. 7, no. 2, 2021.
Authors
Copyright (c) 2026 Journal of Pure & Applied Sciences

This work is licensed under a Creative Commons Attribution 4.0 International License.
In a brief statement, the rights relate to the publication and distribution of research published in the journal of the University of Sebha where authors who have published their articles in the journal of the university of Sebha should how they can use or distribute their articles. They reserve all their rights to the published works, such as (but not limited to) the following rights:
- Copyright and other property rights related to the article, such as patent rights.
- Research published in the journal of the University of Sebha and used in its future works, including lectures and books, the right to reproduce articles for their own purposes, and the right to self-archive their articles.
- The right to enter a separate article, or for a non-exclusive distribution of their article with an acknowledgment of its initial publication in the journal of Sebha University.
Privacy Statement The names and e-mail addresses entered on the Sabha University Journal site will be used for the aforementioned purposes only and for which they were used.