Comparative Analysis of Machine Learning and Deep Learning Models for PM2.5 and PM10 Time Series Forecasting Using the SISANAPAS Web Platform
DOI:
https://doi.org/10.34148/teknika.v14i3.1335Keywords:
Air Quality, Particulate Matter, Time Series Forecasting, Machine Learning, Deep LearningAbstract
Particulate matter (PM2.5 and PM10) pollution remains a significant environmental concern in Indonesia. This study employed the SISANAPAS web platform to compare machine learning (ML) and deep learning (DL) algorithms for PM2.5 and PM10 forecasting. Using 3-hourly data from Cibeureum (January-October 2024), which underwent comprehensive pre-processing (K-Nearest Neighbors imputation, Z-score outlier removal, Min-Max scaling, feature engineering), Long Short-Term Memory (LSTM), Gated Recurrent Unit (GRU), Random Forest, and XGBoost models were evaluated. XGBoost Regression provided the most accurate forecasts, with an R² of 0.85, RMSE of 5.05 µg/m³, and MAE of 3.19 µg/m³ for PM2.5, and an R² of 0.85, RMSE of 9.71 µg/m³, and MAE of 6.55 µg/m³ for PM10. These results, significantly outperforming LSTM and GRU, highlight XGBoost's potential for reliable air quality prediction and demonstrate SISANAPAS as a valuable tool for environmental data analysis, crucial for informing public health and environmental policies.
Downloads
References
[1] A. Rachmawardani, D. Prabowo, Nardi, K. L. Toruan, S. Abdurachman, and D. B. Adinara, “Designing a Cost-Effective IoT Air Quality Sensor for Real-Time Data Monitoring,” 2024 IEEE Int. Conf. Comput. ICOCO 2024, pp. 219–224, 2024, doi: 10.1109/ICOCO62848.2024.10928212.
[2] IQAir, “World Air Quality Report 2023,” IQAir, pp. 1–45, 2023, [Online]. Available: https://www.iqair.com/world-most-polluted-countries
[3] Z. C. Winni and S. Mataram, “Perancangan Zine Penggunaan Transportasi Umum Guna Mengurangi Polusi Udara Bagi Dewasa Di Jakarta,” Tuturrupa, vol. 6, no. 2, pp. 101–110, 2024, doi: 10.24167/tuturrupa.v6i2.11373.
[4] U.S. Environmental Protection Agency, “Particulate Matter (PM) Basics.” Accessed: May 07, 2025. [Online]. Available: https://www.epa.gov/pm-pollution/particulate-matter-pm-basics#PM
[5] U.S. Environmental Protection Agency, “Technical Assistance Document for the Reporting of Daily Air Quality – the Air Quality Index (AQI),” Research Triangle Park, NC, 2024.
[6] H. Kim et al., “Cardiovascular effects of long‐term exposure to air pollution: a population‐based study with 900 845 person‐years of follow‐up,” J. Am. Heart Assoc., vol. 6, no. 11, 2017, doi: 10.1161/jaha.117.007170.
[7] S. M. Daryanoosh, G. Goudarzi, M. J. Mohammadi, H. Armin, Y. O. Khaniabadi, and S. Sadeghi, “Exposure to particulate matter and its health impacts (an airq approach),” Arch. Hyg. Sci., vol. 6, no. 1, pp. 88–95, 2017, doi: 10.29252/archhygsci.6.1.88.
[8] J. D. Sacks et al., “Particulate matter–induced health effects: who is susceptible?,” Environ. Health Perspect., vol. 119, no. 4, pp. 446–454, 2011, doi: 10.1289/ehp.1002255.
[9] J. Li et al., “Major air pollutants and risk of copd exacerbations: a systematic review and meta-analysis,” Int. J. Chron. Obstruct. Pulmon. Dis., vol. Volume 11, pp. 3079–3091, 2016, doi: 10.2147/copd.s122282.
[10] W. Yang, G. Tang, Y. Hao, and J. Wang, “A novel framework for forecasting, evaluation and early-warning for the influence of PM10 on public health,” Atmosphere (Basel)., vol. 12, no. 8, pp. 1–20, 2021, doi: 10.3390/atmos12081020.
[11] P. Sekula, Z. Ustrnul, A. Bokwa, B. Bochenek, and M. Zimnoch, “Random Forests Assessment of the Role of Atmospheric Circulation in PM10 in an Urban Area with Complex Topography,” Sustain., vol. 14, no. 6, pp. 1–43, 2022, doi: 10.3390/su14063388.
[12] A. Pant, S. Sharma, and K. Pant, “Evaluation of Machine Learning Algorithms for Air Quality Index (AQI) Prediction,” J. Reliab. Stat. Stud., vol. 16, no. 2, pp. 229–242, 2023, doi: 10.13052/jrss0974-8024.1621.
[13] Y. T. Tsai, Y. R. Zeng, and Y. S. Chang, “Air pollution forecasting using rnn with lstm,” Proc. - IEEE 16th Int. Conf. Dependable, Auton. Secur. Comput. IEEE 16th Int. Conf. Pervasive Intell. Comput. IEEE 4th Int. Conf. Big Data Intell. Comput. IEEE 3, pp. 1068–1073, 2018, doi: 10.1109/DASC/PiCom/DataCom/CyberSciTec.2018.00178.
[14] R. Amelia, Guskarnali, R. G. Mahardika, C. R. Niani, and N. Lewaherilla, “Predicting particulate matter PM2.5 using the exponential smoothing and Seasonal ARIMA with R studio,” IOP Conf. Ser. Earth Environ. Sci., vol. 1108, no. 1, 2022, doi: 10.1088/1755-1315/1108/1/012079.
[15] N. Zaini, L. W. Ean, A. N. Ahmed, M. Abdul Malek, and M. F. Chow, “PM2.5 forecasting for an urban area based on deep learning and decomposition method,” Sci. Rep., vol. 12, no. 1, pp. 1–13, 2022, doi: 10.1038/s41598-022-21769-1.
[16] Furizal, A. Ma’arif, I. Suwarno, A. Masitha, L. Aulia, and A. N. Sharkawy, “Real-Time Mechanism Based on Deep Learning Approaches for Analyzing the Impact of Future Timestep Forecasts on Actual Air Quality Index of PM10,” Results Eng., vol. 24, no. November, 2024, doi: 10.1016/j.rineng.2024.103434.
[17] H. Altinçöp and A. B. Oktay, “Air Pollution Forecasting with Random Forest Time Series Analysis,” 2018 Int. Conf. Artif. Intell. Data Process. IDAP 2018, pp. 8–12, 2019, doi: 10.1109/IDAP.2018.8620768.
[18] Z. Alfasanah, M. Z. H. Niam, S. Wardiani, M. Ahsan, and M. H. Lee, “Monitoring air quality index with EWMA and individual charts using XGBoost and SVR residuals,” MethodsX, vol. 14, no. August 2024, 2025, doi: 10.1016/j.mex.2024.103107.
[19] YData, “YData Profiling Documentation.” 2024. [Online]. Available: https://docs.profiling.ydata.ai/
[20] IQAir, “Indonesia Air Quality Index (AQI) and Air Pollution Information.” 2025. [Online]. Available: https://www.iqair.com/indonesia
[21] S. Arifin et al., “Long short-term memory (lstm): trends and future research potential,” Int. J. Emerg. Technol. Adv. Eng., vol. 13, no. 5, pp. 24–34, 2023, doi: 10.46338/ijetae0523_04.
[22] A. Sherstinsky, “Fundamentals of Recurrent Neural Network (RNN) and Long Short-Term Memory (LSTM) network,” Phys. D Nonlinear Phenom., vol. 404, no. March, pp. 1–43, 2020, doi: 10.1016/j.physd.2019.132306.
[23] S. Mohsen, “Recognition of human activity using GRU deep learning algorithm,” Multimed. Tools Appl., vol. 82, no. 30, pp. 47733–47749, 2023, doi: 10.1007/s11042-023-15571-y.
[24] Y. Heryadi, Dasar-Dasar Deep Learning dan Implementasinya, I. Yogyakarta: Gava Media, 2021.
[25] A. Rachmawardani, S. K. Wijaya, and A. Shopaheluwakan, “Sistem Peringatan Dini Banjir Berbasis Machine Learning: Studi Literatur,” METHOMIKA J. Manaj. Inform. dan Komputerisasi Akunt., vol. 6, no. 6, pp. 188–198, 2022, doi: 10.46880/jmika.vol6no2.pp188-198.
[26] M. Diaz Resquin et al., “A machine learning approach to address air quality changes during the COVID-19 lockdown in Buenos Aires, Argentina,” Earth Syst. Sci. Data, vol. 15, no. 1, pp. 189–209, 2023, doi: 10.5194/essd-15-189-2023.
[27] C. X. Lv, S. Y. An, B. J. Qiao, and W. Wu, “Time series analysis of hemorrhagic fever with renal syndrome in mainland China by using an XGBoost forecasting model,” BMC Infect. Dis., vol. 21, no. 1, pp. 1–13, 2021, doi: 10.1186/s12879-021-06503-y.
[28] A. Al Mamun, M. Sohel, N. Mohammad, M. S. Haque Sunny, D. R. Dipta, and E. Hossain, “A Comprehensive Review of the Load Forecasting Techniques Using Single and Hybrid Predictive Models,” IEEE Access, vol. 8, pp. 134911–134939, 2020, doi: 10.1109/ACCESS.2020.3010702.
[29] G. E. Saulnier, J. C. Castro, and C. B. Cook, “Forecasting inpatient glycemic control: extension of damped trend methods to subpopulations,” Futur. Sci. OA, vol. 6, no. 10, 2020, doi: 10.2144/fsoa-2020-0096.
[30] T. Esaki, “Appropriate evaluation measurements for regression models,” Chem-Bio Informatics J., vol. 21, no. 0, pp. 59–69, 2021, doi: 10.1273/cbij.21.59.
[31] S. Inayati, N. Iriawan, and I. Irhamah, “A markov switching autoregressive model with time-varying parameters,” Forecasting, vol. 6, no. 3, pp. 568–590, 2024, doi: 10.3390/forecast6030031.
[32] A. F. Sallaby and A. Azlan, “Analysis of Missing Value Imputation Application with K-Nearest Neighbor (K-NN) Algorithm in Dataset,” Int. J. Informatics Comput. Sci., vol. 5, no. 2, pp. 141–144, 2021.
[33] I. J. Fadillah and C. D. Puspita, “Pemanfaatan Metode Weighted K-Nearest Neighbor Imputation (Weighted Knni) Untuk Mengatasi Missing Data,” in Seminar Nasional Official Statistics, 2020, pp. 511–518.
Downloads
Published
Issue
Section
License
Copyright (c) 2025 Teknika

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.















