Implementation of Random Forest Algorithm with Random Oversampling for Sentiment Analysis of X Users Toward the Sekolah Rakyat Program
DOI:
https://doi.org/10.34148/teknika.v15i1.1446Keywords:
Random Forest, Random Oversampling, Sekolah Rakyat Program, Sentiment Analysis, Social Media XAbstract
Social media site X has emerged as a significant platform for voicing public views on government initiatives, such as the Sekolah Rakyat Program. Nevertheless, utilizing social media information for sentiment analysis often faces challenges due to class imbalance, which may result in skewed predictions from models. This research seeks to examine public sentiment and assess how well the Random Forest algorithm performs when paired with Term Frequency–Inverse Document Frequency (TF-IDF) feature extraction and Random Oversampling (ROS) methods to mitigate class imbalance. A dataset comprising 8,623 tweets was gathered and split into training and testing sets using an 80:20 ratio. The results of the experiments indicate that the suggested method demonstrates robust and realistic classification performance, achieving an accuracy of 80.99%, along with a weighted average score in precision, recall, and F1-score of 0.81. Additionally, the sentiment analysis indicates that the majority of public opinions are largely positive, with roughly 69.4% of the testing data reflecting a favorable outlook toward free education access and school improvement efforts. These findings suggest that the proposed model provides reliable performance in capturing public sentiment patterns.
Downloads
References
[1] N. Hadi and D. Sugiarto, “Analisis Sentimen Pembangunan IKN pada Media Sosial X Menggunakan Algoritma SVM, Logistic Regression dan Naïve Bayes,” Jurnal Informatika: Jurnal Pengembangan IT, vol. 10, no. 1, pp. 37–49, Jan. 2025, doi: 10.30591/jpit.v10i1.7106.
[2] W. B. Zulfikar, A. R. Atmadja, and S. F. Pratama, “Sentiment Analysis on Social Media Against Public Policy Using Multinomial Naive Bayes,” Scientific Journal of Informatics, vol. 10, no. 1, pp. 25–34, Jan. 2023, doi: 10.15294/sji.v10i1.39952.
[3] S. N. Russell, L. Rao-Graham, and M. McNaughton, “Mining social media data to inform public health policies: a sentiment analysis case study,” Revista Panamericana de Salud Publica/Pan American Journal of Public Health, vol. 48, 2024, doi: 10.26633/RPSP.2024.79.
[4] “Mensos Lantik 860 Guru Sekolah Rakyat - Tribatanews Polri.”. Dec. 14, 2025. [Online]. Available: https://tribratanews.polri.go.id/blog/nasional-3/mensos-lantik-860-guru-sekolah-rakyat-95697
[5] “Sekolah Rakyat: Gebrakan Satu Tahun Pemerintahan Presiden Prabowo Subianto untuk Memutus Rantai Kemiskinan - Media Keuangan.”. Dec. 14, 2025. [Online]. Available: https://mediakeuangan.kemenkeu.go.id/article/show/sekolah-rakyat-gebrakan-satu-tahun-pemerintahan-presiden-prabowo-subianto-untuk-memutus-rantai-kemiskinan
[6] R. R. Fitri, A. -, and W. Apriandari, “Penggunaan Random Forest Dalam Sistem Klasifikasi Kecemasan Pada Generasi Z,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 13, no. 3, Jul. 2025, doi: 10.23960/jitet.v13i3.6905.
[7] M. S. Kraiem, F. Sánchez-Hernández, and M. N. Moreno-García, “Selecting the suitable resampling strategy for imbalanced data classification regarding dataset properties. An approach based on association models,” Applied Sciences (Switzerland), vol. 11, no. 18, Sep. 2021, doi: 10.3390/app11188546.
[8] T. Wongvorachan, S. He, and O. Bulut, “A Comparison of Undersampling, Oversampling, and SMOTE Methods for Dealing with Imbalanced Classification in Educational Data Mining,” Information (Switzerland), vol. 14, no. 1, Jan. 2023, doi: 10.3390/info14010054.
[9] D. Triyana, M. Muharrom, A. Haromainy, and H. Maulana, “Implementasi Metode Ensemble Majority Vote Pada Algoritma Naive Bayes Dan Random Forest Untuk Analisis Sentimen Twitter Harga Tiket Pesawat Domestik,” Jurnal Mahasiswa Teknik Informatika, vol. 8, no. 4, 2024, doi: 10.36040/jati.v8i4.10475.
[10] H. A. Salman, A. Kalakech, and A. Steiti, “Random Forest Algorithm Overview,” Babylonian Journal of Machine Learning, vol. 2024, pp. 69–79, Dec. 2024, doi: 10.58496/BJML/2024/007.
[11] A. H. Putra and A. Salam, “A Comparative Performance of SMOTE, ADASYN and Random Oversampling in Machine Learning Models on Prostate Cancer Dataset,” Journal of Applied Informatics and Computing (JAIC), vol. 9, no. 3, p. 603, 2025, doi: 10.30871/jaic.v9i3.9308.
[12] S. Diantika, H. Nalatissifa, N. Maulidah, R. Supriyadi, and A. Fauzi, “Penerapan Teknik Random Oversampling Untuk Memprediksi Ketepatan Waktu Lulus Menggunakan Algoritma Random Forest,” Computer Science (CO-SCIENCE), vol. 4, no. 1, pp. 11–18, Jan. 2024, doi: 10.31294/coscience.v4i1.1996.
[13] N. A. Semary, W. Ahmed, K. Amin, P. Pławiak, and M. Hammad, “Enhancing machine learning-based sentiment analysis through feature extraction techniques,” PLoS One, vol. 19, no. 2 February, Feb. 2024, doi: 10.1371/journal.pone.0294968.
[14] A. Anies and M. Ikhsan, “Sentiment Analysis of ‘Free Lunch for Children’ Policy on Social Media X Using Random Forest Algorithm,” Journal of Information Systems and Informatics, vol. 7, no. 1, pp. 649–662, Mar. 2025, doi: 10.51519/journalisi.v7i1.1039.
[15] M. Alif, M. Alamsyah, and M. F. Arif, “Analisis Sentimen Twitter Tentang Pinjaman Online di Indonesia Menggunakan Metode Random Forest,” Jurnal Ilmiah Teknik Informatika Dan Sistem Informasi, vol. 13, no. 2, p. 1410, Aug. 2024, doi: 10.35889/jutisi.v13i2.2215.
[16] S. J. Pipin and H. Kurniawan, “Analisis Sentimen Kebijakan MBKM Berdasarkan Opini Masyarakat di Twitter Menggunakan LSTM,” Jurnal Sifo Mikroskil, vol. 23, pp. 1–5, doi: 10.55601/jsm.v23i2.900.
[17] A. Syah, F. Nurdiyansyah, and A. Y. Rahman, “Analisis Sentimen Aplikasi Shopee, Tokopedia, Lazada dan Blibli Menggunakan Leksikon Dan Random Forest,” Jurnal Informatika dan Teknik Elektro Terapan, vol. 12, no. 3S1, Oct. 2024, doi: 10.23960/jitet.v12i3s1.5155.
[18] X. Zhang, “Text classification algorithms exploration on sentiment analysis,” Applied and Computational Engineering, vol. 5, no. 1, pp. 99–103, May 2023, doi: 10.54254/2755-2721/5/20230541.
[19] W. Xu, J. Chen, Z. Ding, and J. Wang, “Text Sentiment Analysis and Classification Based on Bidirectional Gated Recurrent Units (GRUs) Model.” doi: 10.54254/2755-2721/77/20240670.
[20] A. Apicella, F. Isgrò, and R. Prevete, “Don’t push the button! Exploring data leakage risks in machine learning and transfer learning,” Artif. Intell. Rev., vol. 58, no. 11, Nov. 2025, doi: 10.1007/s10462-025-11326-3.
[21] M. Sivakumar, S. Parthasarathy, and T. Padmapriya, “Trade-off between training and testing ratio in machine learning for medical image processing,” PeerJ Comput. Sci., vol. 10, 2024, doi: 10.7717/PEERJ-CS.2245.
[22] S. Kapoor and A. Narayanan, “Leakage and the Reproducibility Crisis in ML-based Science,” Jul. 2022, [Online]. Available: http://arxiv.org/abs/2207.07048
[23] M. D. Rizkiyanto, M. D. Purbolaksono, and W. Astuti, “Sentiment Analysis Classification on PLN Mobile Application Reviews using Random Forest Method and TF-IDF Feature Extraction,” INTEK: Jurnal Penelitian, vol. 11, no. 1, pp. 37–43, Apr. 2024, doi: 10.31963/intek.v11i1.4774.
[24] G. R. Ashisha, X. A. Mary, E. G. M. Kanaga, J. Andrew, and R. J. Eunice, “Random Oversampling-Based Diabetes Classification via Machine Learning Algorithms,” International Journal of Computational Intelligence Systems, vol. 17, no. 1, Dec. 2024, doi: 10.1007/s44196-024-00678-3.
[25] C. Yang, E. A. Fridgeirsson, J. A. Kors, J. M. Reps, and P. R. Rijnbeek, “Impact of random oversampling and random undersampling on the performance of prediction models developed using observational health data,” J. Big Data, vol. 11, no. 1, Dec. 2024, doi: 10.1186/s40537-023-00857-7.
[26] A. Moreo, A. Esuli, and F. Sebastiani, “Distributional random oversampling for imbalanced text classification,” in SIGIR 2016 - Proceedings of the 39th International ACM SIGIR Conference on Research and Development in Information Retrieval, Association for Computing Machinery, Inc, Jul. 2016, pp. 805–808. doi: 10.1145/2911451.2914722.
[27] M. A. Ganaie, M. Tanveer, P. N. Suganthan, and V. Snasel, “Oblique and rotation double random forest,” Aug. 2022, doi: 10.1016/j.neunet.2022.06.012.
[28] P. Antonio, E. Lim, and C. H. Park, “A Collaborative Ensemble Construction Method for Federated Random Forest,” 2024. doi: 10.1016/j.eswa.2024.124742.
[29] Y. A. Seraphina and P. H. Gunawan, “Sentiment Analysis of Public Responses Regarding The Use of Electric Cars in Indonesia with Support Vector Machine and Random Forest Methods,” The Indonesian Journal of Computer Science, vol. 14, no. 1, Feb. 2025, doi: 10.33022/ijcs.v14i1.4649.
[30] I. A. Nur Sabrina, T. Triastuti Wuryandari, and R. Dapa, “Implementation of Random Forest Algorithm in Classifying Public Sentiment Towards Free Nutritious Meal Program,” Current Science Research Bulletin, vol. 02, no. 05, May 2025, doi: 10.55677/csrb/02-V02I05Y2025.
[31] W. Zhuo and A. Ahmad, “HCRF: an improved random forest algorithm based on hierarchical clustering,” Indonesian Journal of Electrical Engineering and Computer Science, vol. 38, no. 1, p. 578, Apr. 2025, doi: 10.11591/ijeecs.v38.i1.pp578-586.
[32] J. M. Baladjay, N. Riva, L. A. Santos, D. M. Cortez, C. Centeno, and A. A. R. Sison, “Performance evaluation of random forest algorithm for automating classification of mathematics question items,” World Journal of Advanced Research and Reviews, vol. 18, no. 2, pp. 034–043, May 2023, doi: 10.30574/wjarr.2023.18.2.0762.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Teknika

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.















