Optimization of K Value in KNN Algorithm for Spam and HAM Classification in SMS Texts

Authors

DOI:

https://doi.org/10.35870/ijsecs.v4i2.2681

Keywords:

Classification, KNN, SMS Spam

Abstract

Spam refers to the unsolicited and repetitive sending of messages to others via electronic devices without their consent. This activity, commonly known as spamming, is typically carried out by individuals referred to as spammers. SMS spam, which often originates from unknown sources, frequently contains advertisements, phishing attempts, scams, and even malware. Such spam messages can be pervasive, affecting almost all mobile phone numbers, thereby causing significant disruptions to communication by delivering irrelevant content. The persistent nature of spam messages underscores the need for effective filtering mechanisms. This study investigates the application of the K-Nearest Neighbors (KNN) algorithm for classifying SMS messages as either spam or non-spam (ham). The findings demonstrate that KNN, when optimized through various methods for determining the appropriate value of K, can achieve an impressive average accuracy of 99.16% in classifying SMS spam. This high level of accuracy indicates that KNN is a reliable method for spam detection.

Downloads

Download data is not yet available.

Author Biographies

  • Ferryma Arba Apriansyah, Universitas Teknologi Yogyakarta

    Information Technology Study Program-Masters Program, Universitas Teknologi Yogyakarta, Special Region of Yogyakarta, Indonesia

  • Arief Hermawan, Universitas Teknologi Yogyakarta

    Information Technology Study Program-Masters Program, Universitas Teknologi Yogyakarta, Special Region of Yogyakarta, Indonesia

  • Donny Avianto, Universitas Teknologi Yogyakarta

    Information Technology Study Program-Masters Program, Universitas Teknologi Yogyakarta, Special Region of Yogyakarta, Indonesia

References

Nanja, M., & Purwanto, P. (2015). Metode K-Nearest Neighbor berbasis forward selection untuk prediksi harga komoditi lada. Pseudocode, 2(1), 53–64. https://doi.org/10.33369/pseudocode.2.1.53-64

Jain, G., Sharma, M., & Agarwal, B. (2019). Optimizing semantic LSTM for spam detection. International Journal of Information Technology, 11, 239-250. https://doi.org/10.1007/s41870-018-0157-5.

Jindal, N., & Liu, B. (2007, May). Review spam detection. In Proceedings of the 16th international conference on World Wide Web (pp. 1189-1190).

Jiang, M., Cui, P., & Faloutsos, C. (2016). Suspicious behavior detection: Current trends and future directions. IEEE intelligent systems, 31(1), 31-39. https://doi.org/10.1109/MIS.2016.5.

Roul, R. K., Sahoo, J. K., & Arora, K. (2018). Modified TF-IDF term weighting strategies for text categorization. In 2017 14th IEEE India Council International Conference (INDICON) (no. October). https://doi.org/10.1109/INDICON.2017.8487593

Martha, M., Christanti, V., Naga, D. S., & Rompas, P. T. D. (2018). Perbandingan Pengklasifikasi k-Nearest Neighbor dan Neighbor-Weighted k-Nearest Neighbor Pada Sistem Analisis Sentimen dengan Data Microblog. FRONTIERS: JURNAL SAINS DAN TEKNOLOGI, 1(1). https://doi.org/10.36412/frontiers/001035e1/april201801.08

Irfa, A. A., Adiwijaya, A., & Mubarok, M. S. (2018). Klasifikasi Topik Berita Berbahasa Indonesia Menggunakan k-Nearest Neighbor. eProceedings of Engineering, 5(2).

Ling, J., Kencana, I. P. E. N., & Oka, T. B. (2014). Analisis sentimen menggunakan metode Naïve Bayes Classifier dengan seleksi fitur Chi Square. E-Jurnal Matematika, 3(3), 92. https://doi.org/10.24843/mtk.2014.v03.i03.p070

Tamil, N., & Andhra, P. (2020). Classification of social media text spam using VAE-CNN and LSTM mode. Ingénierie des Systèmes d’Information, 25(6), 747-753.

Widyasanti, N. K., Putra, I. D., & Rusjayanthi, N. D. (2018). Seleksi Fitur Bobot Kata dengan Metode TFIDF untuk Ringkasan Bahasa Indonesia. J. Ilm. Merpati (Menara Penelit. Akad. Teknol. Informasi), 6(2), 119.

Zuviyanto, E., Adji, T. B., & Setiawan, N. A. (2018). Perbandingan Algoritme-algoritme Pembelajaran Mesin pada Klasifikasi SMS Spam. Prosiding SENIATI, 4(3), 20-26. https://doi.org/10.36040/seniati.v4i3.1350.

Muzakki, M. A. (2020). Klasifikasi dan Analisa Sentimen Kuesioner Fasilitas dan Layanan untuk Universitas Qomaruddin Gresik. Journal of Computer Science and Visual Communication Design, 5(2), 68-76.

Ramadhan, R., Sari, Y. A., & Adikara, P. P. (2021). Perbandingan Pembobotan Term Frequency-Inverse Document Frequency dan Term Frequency-Relevance Frequency terhadap Fitur N-Gram pada Analisis Sentimen. Jurnal Pengembangan Teknologi Informasi dan Ilmu Komputer, 5(11), 5075-5079.

Herwijayanti, B., Ratnawati, D. E., & Muflikhah, L. (2018). Klasifikasi Berita Online dengan menggunakan Pembobotan TF-IDF dan Cosine Similarity. Jurnal Pengembangan Teknologi Informasi dan Ilmu Komputer, 2(1), 306-312.

Pramartha, G. S., Shaufiah, S., & Bijaksana, M. A. (2015). Analisis Dan Implementasi Algoritma Graph-basedk-nearest Neighbour Untuk Klasifikasi Spam Pada Pesan Singkat. eProceedings of Engineering, 2(2).

Downloads

Published

2024-08-20

How to Cite

Apriansyah, F. A., Hermawan, A., & Avianto, D. (2024). Optimization of K Value in KNN Algorithm for Spam and HAM Classification in SMS Texts. International Journal Software Engineering and Computer Science (IJSECS), 4(2), 767-779. https://doi.org/10.35870/ijsecs.v4i2.2681

Similar Articles

16-20 of 30

You may also start an advanced similarity search for this article.

Most read articles by the same author(s)