Implementation of Apache Airflow for Pipeline Automation: Case Study of Mental Health Issues on Social Media X
DOI:
https://doi.org/10.34148/teknika.v15i2.1473Keywords:
ETL Pipeline, Apache Airflow, Data Engineering, Mental Health, DockerAbstract
The increasing use of social media as a platform for discussing mental health issues presents significant opportunities for psychological and social media research. However, the lack of automated, fail-resilient, and clean data provisioning remains a major infrastructure challenge. This study developed an automated ETL (Extract, Transform, Load) pipeline using Apache Airflow and Docker Compose to collect and process bilingual social media data from the X platform (Twitter). The system integrates a container-based architecture with persistent volumes to ensure infrastructure resilience and data continuity. A bilingual text preprocessing pipeline was implemented to support automated cleaning and normalization of English and Indonesian social media content. Experimental results show that the system successfully collected 1,768 unique tweets with a preprocessing throughput of 75 tweets per second and achieved a 39% reduction in average text length as part of preprocessing efficiency measurement. These measurements are intended to evaluate ETL preprocessing efficiency and throughput rather than analytical or linguistic quality. Resilience testing confirmed 100% data integrity with zero data loss across multiple failure scenarios, including container restarts and process suspensions. In addition, the implementation of ShortCircuitOperator successfully mitigated API credit exhaustion (HTTP 402) through a graceful degradation mechanism. This study contributes a reproducible and reliable data engineering framework for automated social media data collection and preprocessing, resulting in structured bilingual datasets ready for downstream analytical tasks. However, mental health in this study is positioned as a domain-specific case study, and the resulting dataset is not intended to serve as a clinical diagnostic tool.
Downloads
References
[1] M. K. D. Jaya, I. G. A. G. A. Kadnyanan, and D. L. G. Astuti, “ETL Pipeline Di Pt Sirkadian Indonesia Sehat,” Jurnal Pengabdian Informatika, vol. 4, no. 2, pp. 589–594, 2024.
[2] Z. Aurellia, E. Sahara, M. Ilham, Z. Welerubun, and O. D. Ramadhani, “Pengalaman Pengguna X Sebagai Platform Curhat (Studi Fenomenologi),” Seminar Nasional Universitas Negeri Surabaya, vol. 3, pp. 616–623, 2024.
[3] Kukuh Wijayanti and Qoniah Nur Wijayani, “Peranan Aplikasi Twitter (X) dalam Interaksi Komunikasi untuk Membantu Menyeimbangkan Kesehatan Mental pada Remaja Saat Ini,” Jurnal Sains Student Research, vol. 2, no. 1, pp. 07–15, Dec. 2023, doi: 10.61722/jssr.v2i1.469.
[4] D. K. N. Manullang, F. Veronica, E. Anastasia, M. Akailupa, D. Chrisnawandi, and R. Gunawan, “Dampak Penggunaan Twitter Mempengaruhi Kesehatan Mental Generasi Z,” TERAPUTIK: Jurnal Bimbingan dan Konseling, vol. 8, no. 3, pp. 91–103, Feb. 2025, doi: 10.26539/teraputik.833609.
[5] M. Rahman Nayla, “Memahami Dampak Media Sosial terhadap Kesehatan Mental Mahasiswa,” Jimad: Jurnal Ilmiah Mutiara Pendidikan, vol. 2, no. 1, 2024, doi: 10.31004/jptam.v7i3.11993.
[6] Rizqa Annisa Dwiatmoko and Imelda, “Analisis Sentimen Kesehatan Mental Di Twitter Menggunakan Naive Bayes Classifier Dan K-Nearest Neighbor,” JIFOSI, vol. 6, no. 1, pp. 1–9, Apr. 2025, doi: 10.33005/jifosi.v6i1.466.
[7] M. R. Hidayatullah and W. Maharani, “Depression Detection on Twitter Social Media Using Decision Tree,” Jurnal Resti, vol. 6, no. 4, pp. 677–683, Aug. 2022, doi: 10.29207/resti.v6i4.4275.
[8] H. Ramadhan, Y. Hendriyani, and T. Sriwahyuni, “Penerapan Pipeline ETL Apache Airflow Menggunakan Algoritma Collaborative Filtering untuk Pemberian Rekomendasi Gastrodiplomasi Indonesia,” Jurnal Pendidikan Tambusai, vol. 9, no. 2, 2025, doi: https://doi.org/10.31004/jptam.v9i2.27485.
[9] L. Fanani Mz et al., “Rekontruksi Arsitektur DataBase untuk Peningkatan Proses Load Data,” Jurnal Media Informatika (JUMIN), vol. 6, no. 2, pp. 1455–1460, 2025, doi: https://doi.org/10.55338/jumin.v6i2.5585.
[10] D. Andriansyah, “Implementasi Extract-Transform-Load (ETL) Data Warehouse Laporan Harian Pool,” Jurnal Teknik Informatika Stimik Antar Bangsa, vol. VIII, no. 2, pp. 45–49, Dec. 2022, doi: https://doi.org/10.51998/jti.v8i2.486.
[11] S. Kumar Singu, “ETL Process Automation: Tools and Techniques,” ESP Journal of Engineering & Technology Advancements, vol. II, no. 1, pp. 74–85, Feb. 2022, doi: 10.56472/25832646/JETA-V2I1P110.
[12] D. Chanda, “Automated ETL Pipelines for Modern Data Warehousing: Architectures, Challenges, and Emerging Solutions,” The Eastasouth Journal of Information System and Computer Science, vol. 1, no. 03, pp. 209–212, 2024, doi: 10.58812/esiscs.v1i03.
[13] U. Nayak, “Comparing AWS Glue vs. Apache Airflow for Data Orchestration: A Comprehensive Performance and Cost Analysis,” International Journal of Emerging Trends in Computer Science and Information Technology, vol. 6, no. 3, pp. 51–55, 2025, doi: 10.63282/3050-9246.ijetcsit-v6i3p109.
[14] E. Eko Wahyudi et al., “Akuisisi Data Prediksi Curah Hujan Secara Periodik Menggunakan Apache Airflow,” Badan Meteorologi, Klimatologi, dan Geofisika Jl. Angkasa I, vol. 4, no. 2, pp. 1–012, doi: 10.20895/INISTA.V4I2.
[15] I. R. Wibowo et al., “Big Data Pipeline Infrastructure Design in MSME E-Commerce Systems with a Focus on Data Source Processing Using Orchestration Tools,” Transmisi: Jurnal Ilmiah Teknik Elektro, vol. 26, no. 1, pp. 48–54, Jan. 2024, doi: 10.14710/transmisi.26.1.48-54.
[16] S. Harsha and V. Sanne, “Addressing Persistent Storage Challenges in Kubernetes Environments,” International Journal on Recent and Innovation Trends in Computing and Communication, [Online]. Available: http://www.ijritcc.org
[17] Z. Bakhshi, G. Rodriguez-Navas, and H. Hansson, “Analyzing the performance of persistent storage for fault-tolerant stateful fog applications,” Journal of Systems Architecture, vol. 144, Nov. 2023, doi: 10.1016/j.sysarc.2023.103004.
[18] J. H. Na et al., “PVA: The Persistent Volume Autoscaler for Stateful Applications in Kubernetes,” IEEE Access, vol. 12, pp. 179130–179143, 2024, doi: 10.1109/ACCESS.2024.3507194.
[19] M. Musrini, N. Fitrianti, F. Reysandi, A. Rifky, N. Fitrianti F, and R. A. Rifky, “Membangun Data Warehouse untuk Menganalisis Pola Siswa yang Mendaftar di ITENAS (Studi Kasus Institut Teknologi Nasional Bandung),” Jurnal Ilmiah Teknologi Informasi Terapan, vol. 8, no. 1, pp. 45–56, Dec. 2021, doi: https://doi.org/10.33197/jitter.vol8.iss1.2021.715.
[20] A. W. Iswara, H. Setiadi, and A. Wijayanto, “Implementation of Business Intelligence for Quality Support of RSUD Ir. Soekarno Sukoharjo with Data Warehouse,” ITSMART: Jurnal Ilmiah Teknologi dan Informasi, vol. 9, no. 1, Jun. 2020, doi: https://doi.org/10.20961/itsmart.v9i1.42915.
[21] I. Jubaidah, D. Pratiwi, and T. Siswanto, “Forecasting sales data on e-commerce using single exponential smoothing methods,” Jurnal Teknologi Informasi : Jurnal Keilmuan dan Aplikasi Bidang Teknik Informatika, vol. 17, no. 2, Aug. 2023, doi: 10.47111/JTI.
[22] I. Putu, W. Prasetia, I. Nyoman, and H. Kurniawan, “Implementasi ETL (Extract, Transform, Load) pada Data warehouse Penjualan Menggunakan Tools Pentaho,” TIERS Information Technology Journal, vol. 2, no. 1, pp. 39–47, 2021.
[23] D. Wijayanto and A. Firdonsyah, “Implementation of Continous Delivery using Jenkins And Kubernetes with Docker Local Images,” sinkron, vol. 8, no. 4, pp. 2226–2235, Oct. 2023, doi: 10.33395/sinkron.v8i4.12624.
[24] N. Reddy Rachamala, “Automated Orchestration for Distributed ETL: Practical Approaches with Airflow,” Journal of Information Systems Engineering and Management, vol. 2025, no. 59s, pp. 2468–4376, 2025.
[25] A. Satya Vivek Vardhan Akisetty et al., “Automating ETL Workflows with CI/CD Pipelines for Machine Learning Applications,” Iconic Research and Engineering Journals (IRE Journals), vol. 7, no. 3, Feb. 2025.
[26] R. Tahir, G. Surya Mahendra, and R. Sandra Yofa Zebua, Bisnis Intelegent (Pengantar Business Intelligence dalam Bisnis). Cilegon: PT. Sonpedia Publishing Indonesia, 2023.
[27] R. W. P. G. T. Kusumah, “Perancangan Data Warehouse Pada Bagian Akademik Universitas di Bandung,” Jurnal Strategi (Sarana Tugas Akhir Mahasiswa Teknologi Informasi), vol. 3, no. 1, May 2021.
[28] I. Zaelani, “Implementasi Data Mart Terhadap Sistem Penjualan Pada Perusahaan Bidang Distributor Di Pt. Eigen Trimathema Implementation Of Data Mart On Sales System In Distributor Companies In Pt. Eigen Trimathema,” JUPITER: Jurnal Penelitian Mahasiswa Teknik Dan Ilmu Komputer, vol. 1, no. 2, 2021.
[29] I. Gusti Ngurah Agung Trisna Putra et al., “Implementasi ETL Data Warehouse Dengan Konsep Fitur Metadata Dan Cleansing Data Pada Toko Kue,” Jurnal Sistem Informasi, vol. 9, no. 2, pp. 274–289, 2020.
[30] K. Rahayu, V. Fitria, D. Septhya, R. Rahmaddeni, and L. Efrizoni, “Klasifikasi Teks untuk Mendeteksi Depresi dan Kecemasan pada Pengguna Twitter Berbasis Machine Learning,” MALCOM: Indonesian Journal of Machine Learning and Computer Science, vol. 3, no. 2, pp. 108–114, Sep. 2023, doi: 10.57152/malcom.v3i2.780.
[31] S. Shevira, I. Made, A. Dwi Suarjaya, W. Buana, J. Raya, and K. Udayana, “Lexicon and Naive Bayes Algorithms to Detect Mental Health Situations from Twitter Data,” Journal of Information Systems Engineering and Business Intelligence, vol. 8, no. 2, 2022, doi: 10.20473/jisebi.8.2.
[32] Nithish, Ravi, and David, “Data Transformation Techniques in ETL,” International Journal of Multidisciplinary on Science and Management, vol. 1, pp. 1–16, doi: 10.71141/30485037/V1I2P101.
[33] I. Zaelani, “Implementasi Data Mart terhadap Sistem Penjualan pada Perusahaan Distributor di PT. Eigen Trimathema,” Jupiter: Jurnal Penelitian Mahasiswa Teknik Dan Ilmu Komputer, vol. 1, no. 2, 2021.
[34] R. Wijaya and B. Pudjoatmodjo, “Penerapan Extraction-Transformation-Loading (ETL) Dalam Data Warehouse (Studi Kasus: Departemen Pertanian),” Jurnal Nasional Pendidikan Teknik Informatika (JANAPATI), vol. 5, no. 2, 2016.
[35] Nithish, Ravi, and David, “Data Transformation Techniques in ETL,” International Journal of Multidisciplinary on Science and Management, vol. 1, pp. 1–16, doi: 10.71141/30485037/V1I2P101.
[36] J. Lee et al., “MDB-KCP: persistence framework of in-memory database with CRIU-based container checkpoint in Kubernetes,” Journal of Cloud Computing, vol. 13, no. 1, Dec. 2024, doi: 10.1186/s13677-024-00687-9.
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Teknika

This work is licensed under a Creative Commons Attribution-ShareAlike 4.0 International License.















