Multi-Platform Sentiment Analysis of Diabetes Mellitus on X and TikTok Using K-Nearest Neighbor, Chi-Square Selection, and Oversampling
DOI:
https://doi.org/10.35870/ijmsit.v6i2.8061Keywords:
Sentiment Analysis, Diabetes Mellitus, K-Nearest Neighbor, Chi-Square, SMOTEAbstract
Diabetes mellitus is a chronic health condition that is widely discussed by the public on social media, generating a large volume of opinions that are difficult to interpret manually. This study analyzes public sentiment toward Diabetes mellitus using data collected from X (Twitter) and TikTok. Text data were preprocessed (cleaning, slang normalization, stopword removal, and stemming) and duplicate entries were removed. Sentiment labels were generated automatically through a rule-based lexicon across four categories (positive, negative, neutral, and irrelevant/discarded), and the reliability of this automatic labeling was verified through manual validation of a stratified sample of 200 data points, measured using Cohen's Kappa. Positive and negative data were then weighted using TF-IDF, reduced using Chi-Square feature selection, balanced using SMOTE, and classified using K-Nearest Neighbor (KNN). Model performance was evaluated using accuracy, precision, recall, and F1-score, and compared across four scenarios: baseline KNN, KNN with Chi-Square, KNN with SMOTE, and the combined KNN+Chi-Square+SMOTE model. The manual validation of 198 valid samples produced an agreement accuracy of 59.60% and a Cohen's Kappa of 0.459 (moderate agreement), indicating that the main source of disagreement lies at the boundary between the neutral and sentiment-bearing classes, while direct positive-negative misclassification was rare (4.5%). The combined model achieved an accuracy of 82.86%, with a macro-averaged precision of 82.95%, recall of 83.56%, and F1-score of 82.79%. Interestingly, Chi-Square feature selection alone yielded the highest accuracy among the four scenarios (85.10%), suggesting that feature selection contributed more to performance gains than class balancing in this dataset. These findings suggest that combining feature selection and oversampling techniques improves the reliability of multi-platform sentiment classification for health-related topics and can inform more effective public health communication strategies regarding diabetes.
Downloads
References
Aminuddin, A., Sima, Y., Izza, N. C., Lalla, N. S. N., & Arda, D. (2023). Edukasi Kesehatan Tentang Penyakit Diabetes Melitus bagi Masyarakat. Abdimas Polsaka, 7–12. https://doi.org/10.35816/abdimaspolsaka.v2i1.25
Asro'i, A., & Februariyanti, H. (2022). Analisis Sentimen Pengguna Twitter Terhadap Perpanjangan PPKM Menggunakan Metode K-Nearest Neighbor. Jurnal Khatulistiwa Informatika, 10(1), 17–24. https://doi.org/10.31294/jki.v10i1.12624
Bhuana, K., Indriati, & Muflikhah, L. (2022). Analisis Sentimen Masyarakat Indonesia tentang Vaksin Covid-19 di Twitter dengan menggunakan Metode K-Nearest Neighbors dan Seleksi Fitur Chi Square. Jurnal Pengembangan Teknologi Informasi dan Ilmu Komputer, 6(3), 1395–1401.
Candra, C., Chandra, K. W., & Irsyad, H. (2024). Efektifitas SMOTE dalam Mengatasi Imbalanced Class Algoritma K-Nearest Neighbors pada Analisis Sentimen terhadap Starlink. Jurnal Ilmu Komputer dan Informatika, 4(1), 31–42. https://doi.org/10.54082/jiki.132
Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33(1), 159–174. https://doi.org/10.2307/2529310
Pramayasa, K., Maysanjaya, I. M. D., & Indradewi, I. G. A. A. D. (2023). Analisis Sentimen Program MBKM Pada Media Sosial Twitter Menggunakan KNN dan SMOTE. SINTECH (Science and Information Technology) Journal, 6(2), 89–98. https://doi.org/10.31598/sintechjournal.v6i2.1372
Qadri, M. (2020). Pengaruh Media Sosial Dalam Membangun Opini Publik. Qaumiyyah: Jurnal Hukum Tata Negara, 1(1), 49–63. https://doi.org/10.24239/qaumiyyah.v1i1.4
Setiawan, I., & Andriyani, W. (2026). Perbandingan Kinerja dan Efisiensi Model NLP pada Analisis Sentimen Ulasan Aplikasi Layanan Publik Digital. Jurnal Sistem Komputer dan Informatika (JSON), 7(4), 1505–1517. https://doi.org/10.30865/json.v7i4.9774
Shefia, F. A., Setiaji, P., & Triyanto, W. A. (2026). Analisis Sentimen Ulasan Aplikasi CapCut pada Google Play Store menggunakan Support Vector Machine dengan Teknik SMOTE. Sistemasi: Jurnal Sistem Informasi, 15, 658–668. https://doi.org/10.32520/stmsi.v15i2.5948
Ubaidillah, M., Fatah, D. A., & Negara, Y. D. P. (2025). Penerapan SMOTE dan Chi-Square Feature Selection untuk Meningkatkan Akurasi Model Multinomial Naïve Bayes dalam Analisis Sentimen Video "Presiden Prabowo Menjawab" di YouTube. Jurnal Nasional Komputasi dan Teknologi Informasi (JNKTI), 8(5). https://doi.org/10.32672/jnkti.v8i5.9807
Vidya Sakta, P., Indriati, & Marji. (2020). Analisis Sentimen Pariwisata di Kabupaten Malang dengan Menggunakan Metode BM25F, Neighbor Weighted K-Nearest Neighbor dan Seleksi Fitur Chi-Square. Jurnal Pengembangan Teknologi Informasi dan Ilmu Komputer, 4(10), 3659–3666.
Wahyuni, S., Arisani, G., Riani, R., & Hanipah, H. (2022). Peran Media Sosial Sebagai Upaya Promosi Kesehatan. Jurnal Forum Kesehatan: Media Publikasi Kesehatan Ilmiah, 11(2), 86–96. https://doi.org/10.52263/jfk.v11i2.233
Wijaya, R., & Suwandhi, A. (2024). Sentimen Komentar Universitas Pelita Harapan Pada TikTok Menggunakan Metode K-Nearest Neighbor. JDMIS: Journal of Data Mining and Information Systems, 2(1), 26–36. https://doi.org/10.54259/jdmis.v2i1.2418
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Asti Nisa, Noor Latifah, R. Rhoedy Setiawan

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
1. Copyright Retention and Open Access License
Authors retain copyright of their work and grant the journal non-exclusive right of first publication under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
This license allows unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
2. Rights Granted Under CC BY 4.0
Under this license, readers are free to:
- Share — copy and redistribute the material in any medium or format
- Adapt — remix, transform, and build upon the material for any purpose, including commercial use
- No additional restrictions — the licensor cannot revoke these freedoms as long as license terms are followed
3. Attribution Requirements
All uses must include:
- Proper citation of the original work
- Link to the Creative Commons license
- Indication if changes were made to the original work
- No suggestion that the licensor endorses the user or their use
4. Additional Distribution Rights
Authors may:
- Deposit the published version in institutional repositories
- Share through academic social networks
- Include in books, monographs, or other publications
- Post on personal or institutional websites
Requirement: All additional distributions must maintain the CC BY 4.0 license and proper attribution.
5. Self-Archiving and Pre-Print Sharing
Authors are encouraged to:
- Share pre-prints and post-prints online
- Deposit in subject-specific repositories (e.g., arXiv, bioRxiv)
- Engage in scholarly communication throughout the publication process
6. Open Access Commitment
This journal provides immediate open access to all content, supporting the global exchange of knowledge without financial, legal, or technical barriers.
