Pemodelan Hybrid untuk Prediksi Risiko Keparahan Penyakit Tuberkulosis Menggunakan Algoritma K-Means dan Random Forest
DOI:
https://doi.org/10.35870/jtik.v10i3.6379Keywords:
Tuberculosis, K-Means, Random Forest, Hybrid ModelAbstract
Tuberculosis (TB) remains a major infectious disease in Indonesia, while the identification of patient severity levels in healthcare facilities is often time-consuming due to manual assessment of medical records. At Puskesmas Bonang 1, TB cases increased from 41 in 2023 to 57 in 2024, yet no data-driven analytical system is available to support rapid and objective risk evaluation. This study utilizes 2,546 TB patient medical records from 2023–2024 and applies preprocessing, normalization, encoding, clustering using K-Means, and the development of both baseline and hybrid models. The evaluation results indicate that the Hybrid K-Means + Random Forest model with hyperparameter tuning outperforms the standalone Random Forest model. The baseline Random Forest achieved an accuracy of 81.72% with an F1-Score of 80.98%, while the Hybrid + Tuning model obtained an accuracy of 82.51% and an F1-Score of 81.34%. This improvement demonstrates that cluster-based features extracted using K-Means successfully enhance data representation and improve the predictive performance of Tuberculosis severity risk classification.
Downloads
References
Agusyul, A. Y., & Firmansyah, F. (2023). Prediksi Penyakit Jantung Menggunakan Algoritma Random Forest. Jurnal Minfo Polgan, 12(2). https://doi.org/10.33395/jmp.v12i2.13214.
Airlangga, G. (2024). Leveraging Machine Learning for Accurate Anemia Diagnosis Using Complete Blood Count Data. Indonesian Journal of Artificial Intelligence and Data Mining, 7(2), 318. https://doi.org/10.24014/ijaidm.v7i2.29869.
Arif, M., Setiawan, M., Dwi Hartono, A., Arif Ma, M., & Setiawan, R. (2025). Menggunakan Metode Machine Learning Untuk Memprediksi Nilai Mahasiswa Dengan Model Prediksi Multiclass. Jurnal Informatika: Jurnal Pengembangan IT, 10(1).
Bagus Pratama, Y., & Setiawan, A. (2024). RESOLUSI: Rekayasa Teknik Informatika dan Informasi Implementasi Machine Learning Menggunakan Algoritma K-Means Untuk Klasifikasi Sekolah Dasar. Media Online, 4(3).
Helmiyah, S., Pramestiawan, R., Studi Pendidikan Informatika, P., & Tinggi Keguruan dan Ilmu Pendidikan Rosalia Lampung, S. (2025). Analisis Komparatif Algoritma Machine Learning dengan Metrik Akurasi, Presisi, Recall, dan F1-Score pada Dataset Kacang Kering. IKOMTI, 6(3), 152–159. https://doi.org/10.35960/ikomti.v6i3.2031.
Juliasih, N. N., Soedarsono, & Sari, R. M. (2020). Analysis of tuberculosis program management in primary health care. Infectious Disease Reports, 12. https://doi.org/10.4081/idr.2020.8728.
Kevin, J., & Wijaya, B. A. (2025). Implementasi Algoritma Clustering dan Classification dalam Data Mining: Systematic Literature Review terhadap Tren dan Tantangan Terkini. Jurnal Publikasi Sistem Informasi Dan Manajemen Bisnis, 4(3), 372–386. https://doi.org/10.55606/jupsim.v4i3.5240.
Laksono, B., Syahidin, Y., & Yunengsih, Y. (2024). Implementasi Data Mining Klasterisasi Data Pasien Rawat Inap dengan Algoritma K-Means Clustering. Jurnal Teknologi Sistem Informasi Dan Aplikasi, 7(2), 621–627. https://doi.org/10.32493/jtsi.v7i2.39354.
Mutawali, L., Murniati, W., & Kunci, K. (2022). Penerapan KNNImputer dalam mengolah data missing value untuk membantu meningkatkan akurasi Support Vector Machine klasifikasi penyakit tiroid. JINTEKS, 4(4).
Nasution, F. A., & Juledi, A. P. (2025). Penerapan Algoritma Random Forest untuk Klasifikasi Tingkat Keparahan Penyakit pada Data Rekam Medis. Journal of Computer Science and Information Systems, 371–378.
Nugroho, A. K., Hayati, L. N., & Jabir, S. R. (2025). Analisis Perbandingan Metode Naive Bayes dan Random Forest pada Klasifikasi Sentimen Publik terhadap Aplikasi Identitas Kependudukan Digital (IKD). Jurnal Algoritma, 22(2). https://doi.org/10.33364/algoritma/v.22-2.2729.
Pratiwi, A., Manurung, H., Pita Uli Sitompul, M., Informasi, S., & Kaputama, S. (2025). Implementasi metode K-Means Clustering untuk mengklasifikasi status gizi balita berdasarkan wilayah kerja Puskesmas Karang Rejo. Great Journal.
Rahman, F. I., Lukman, & Hildayanti. (2025). Arus Jurnal Sains dan Teknologi (AJST) Sistem Klasifikasi Kerusakan Jalan Metode Machine Learning dengan Algoritma K-Means dan Random Forest. AJST, 3(1).
Ramadhani, F., Septiana, D., Amalia, S. N., Fadilah, P. M., & Satria, A. (2024). Klasifikasi risiko gizi buruk pada ibu hamil menggunakan metode Random Forest. Jurnal Teknologi Informasi, 5(2). https://doi.org/10.46576/djtechno.
Renny Afriany, Rudolf Sinaga, & Samsinar Samsinar. (2025). Segmentasi Pasien Berbasis K-Means dari Tanda-tanda Vital dan Demografi: Pendekatan Unsupervised Learning untuk Profil Risiko Klinis. Jurnal Ilmiah Sistem Informasi Dan Ilmu Komputer, 5(3), 638–650. https://doi.org/10.55606/juisik.v5i3.1811.
Riansyah, M., Bahri, S., & Fiqna, H. P. (2025). Perbandingan Akurasi Algoritma Decision Tree dengan Random Forest dalam Diagnosa Penyakit Hepatitis. Indonesian Journal of Science, 2(3).
Setiawan, A., Hadryan Nst, Z., Khairi, Z., & Efrizoni, L. (2024). Klasifikasi tingkat risiko diabetes menggunakan algoritma Random Forest. Jurnal Informatika & Rekayasa Elektronika (JIRE), 7(2).
Susanto, E. R., Inzaghi, M. R., Amarudin, A., & Neneng, N. (2025). Evaluasi Kinerja Model Random Forest Dalam Memprediksi Diabetes Berdasarkan Dataset Kesehatan di Indonesia. Jurnal Pendidikan Dan Teknologi Indonesia, 5(7), 1857–1866. https://doi.org/10.52436/1.jpti.871.
Umar, M., Adytia, P., Yusnita, A., Widya Cipta Dharma, S., Yamin No, J. M., Kelua, G., Samarinda Ulu, K., Samarinda, K., & Timur, K. (2025). Penerapan Algoritma Support Vector Machine (SVM) dalam Analisis Sentimen Mahasiswa Terhadap Sistem Layanan KPST STMIK Widya Cipta Dharma. Jurnal Nasional Komputasi Dan Teknologi Informasi (JNKTI), 8(3).
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Hasan Ibrohim, Harminto Mulyo, Gentur Wahyu Nyipto Wibowo

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
1. Copyright Retention and Open Access License
Authors retain copyright of their work and grant the journal non-exclusive right of first publication under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
This license allows unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
2. Rights Granted Under CC BY 4.0
Under this license, readers are free to:
- Share — copy and redistribute the material in any medium or format
- Adapt — remix, transform, and build upon the material for any purpose, including commercial use
- No additional restrictions — the licensor cannot revoke these freedoms as long as license terms are followed
3. Attribution Requirements
All uses must include:
- Proper citation of the original work
- Link to the Creative Commons license
- Indication if changes were made to the original work
- No suggestion that the licensor endorses the user or their use
4. Additional Distribution Rights
Authors may:
- Deposit the published version in institutional repositories
- Share through academic social networks
- Include in books, monographs, or other publications
- Post on personal or institutional websites
Requirement: All additional distributions must maintain the CC BY 4.0 license and proper attribution.
5. Self-Archiving and Pre-Print Sharing
Authors are encouraged to:
- Share pre-prints and post-prints online
- Deposit in subject-specific repositories (e.g., arXiv, bioRxiv)
- Engage in scholarly communication throughout the publication process
6. Open Access Commitment
This journal provides immediate open access to all content, supporting the global exchange of knowledge without financial, legal, or technical barriers.
