Developing an Indonesian Fake News Detection System Using IndoBERT and ProtoNet for Few-Shot Learning

Authors

  • Yefta Christian Universitas International Batam
  • Wilson Wilson Universitas International Batam
  • Andik Yulianto Universitas International Batam

DOI:

https://doi.org/10.35870/ijsecs.v6i3.8156

Keywords:

Fake News Detection, IndoBERT, Prototypical Network, Few-Shot Learning, Explainable Artificial Intelligence

Abstract

Fake news dissemination through digital media has become a critical concern because inaccurate information may spread swiftly and influence public comprehension, social behavior, and decision-making. This study developed an Indonesian fake news detection web application employing fine-tuned IndoBERT and an IndoBERT-based Prototypical Network (ProtoNet) under few-shot learning scenarios. The research followed the Cross Industry Standard Process for Data Mining framework, consisting of business understanding, data understanding, data preparation, modeling, evaluation, and deployment. The dataset was obtained from the publicly available Deteksi Berita Hoaks Indo Dataset on Kaggle and contained 23,944 Indonesian news records, consisting of 11,200 factual and 12,744 hoax articles. Three few-shot scenarios were evaluated: 50-shot, 100-shot, and 300-shot per class. IndoBERT was fine-tuned as a supervised classification baseline, while ProtoNet used IndoBERT embeddings to construct class prototypes and classify texts based on Euclidean distance. The best performance was achieved by fine-tuned IndoBERT in the 300-shot scenario, with 97.12% accuracy and a 97.10% macro F1-score. ProtoNet achieved 92.65% accuracy and a 92.63% macro F1-score in the same scenario. Even in the most constrained 50-shot setting, ProtoNet achieved an F1-score of 90.82%, closely trailing fine-tuned IndoBERT at 91.17% and demonstrating strong metric stability. The web application supports manual text input, URL-based article extraction, model selection, confidence scores, and word-level occlusion explanations. These results demonstrate the effectiveness of IndoBERT for Indonesian fake news detection and the feasibility of prototype-based classification under limited-data conditions. Overall, this work bridges the gap between deep contextual representation, data-efficient learning, and practical, interpretable deployment for public media verification.

Downloads

Download data is not yet available.

Author Biographies

  • Yefta Christian, Universitas International Batam

    Universitas International Batam, Batam City, Riau Islands Province, Indonesia.

  • Wilson Wilson, Universitas International Batam

    Universitas International Batam, Batam City, Riau Islands Province, Indonesia.

  • Andik Yulianto, Universitas International Batam

    Universitas International Batam, Batam City, Riau Islands Province, Indonesia.

References

Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115. https://doi.org/10.1016/j.inffus.2019.12.012

Azizah, S. F. N., Cahyono, H. D., Sihwi, S. W., & Widiarto, W. (2023). Performance analysis of transformer based models (BERT, ALBERT and RoBERTa) in fake news detection. In 2023 6th International Conference on Information and Communications Technology (ICOIACT) (pp. 425–430). IEEE. https://doi.org/10.1109/ICOIACT59844.2023.10455849

Bao, Y., Wu, M., Chang, S., & Barzilay, R. (2019). Few-shot text classification with distributional signatures. arXiv preprint arXiv:1908.06039. https://doi.org/10.48550/arXiv.1908.06039

Christian, Y., & Yap Rui Qi, K. O. (2022). Penerapan K-Means pada segmentasi pasar untuk riset pemasaran pada startup early stage dengan menggunakan CRISP-DM. JURIKOM (Jurnal Riset Komputer), 9(4), 966–973. https://doi.org/10.30865/jurikom.v9i4.4486

Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423

Dopierre, T., Gravier, C., & Logerais, W. (2021). A neural few-shot text classification reality check. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 935–943). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.eacl-main.79

Erlina, E., Julyanto, J., Rustandi, J., Alexander, A., Francisco, L., Ma’muriyah, N., & Sabariman, S. (2023). Penerapan artificial intelligence pada aplikasi chatbot sebagai sistem pelayanan dan informasi online pada sekolah. Journal of Information System and Technology, 4(3), 610–621. https://doi.org/10.37253/joint.v4i3.6296

Fakhruzzaman, M. N., & Gunawan, S. W. (2021). Web-based application for detecting Indonesian clickbait headlines using IndoBERT. arXiv preprint arXiv:2102.10601. https://doi.org/10.48550/arXiv.2102.10601

Fathin, M. A., Sibaroni, Y., & Prasetyowati, S. S. (2024). Handling imbalance dataset on hoax Indonesian political news classification using IndoBERT and random sampling. Jurnal Media Informatika Budidarma, 8(1), 352–360. https://doi.org/10.30865/mib.v8i1.7099

Fawaid, J., Awalina, A., Krisnabayu, R. Y., & Yudistira, N. (2021). Indonesia’s fake news detection using transformer network. In Proceedings of the 6th International Conference on Sustainable Information Engineering and Technology (pp. 247–251). Association for Computing Machinery. https://doi.org/10.1145/3479645.3479666

Geng, R., Li, B., Li, Y., Zhu, X., Jian, P., & Sun, J. (2019). Induction networks for few-shot text classification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (pp. 3904–3913). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1403

Han, L., Zhang, X., Zhou, Z., & Liu, Y. (2024). A multifaceted reasoning network for explainable fake news detection. Information Processing & Management, 61(6), Article 103822. https://doi.org/10.1016/j.ipm.2024.103822

Jocelynne, C., Wijayakusuma, I. G. N. L., & Harini, L. P. I. (2025). Detection of political hoax news using fine-tuning IndoBERT. Journal of Applied Informatics and Computing, 9(2), 354–360. https://doi.org/10.30871/jaic.v9i2.8989

Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. In Proceedings of the 28th International Conference on Computational Linguistics (pp. 757–770). International Committee on Computational Linguistics. https://doi.org/10.18653/v1/2020.coling-main.66

Lei, S., Zhang, X., He, J., Chen, F., & Lu, C.-T. (2023). TART: Improved few-shot text classification using task-adaptive reference transformation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 11014–11026). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.617

Liu, H., Wang, W., Li, H., & Li, H. (2024). TELLER: A trustworthy framework for explainable, generalizable and controllable fake news detection. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 15556–15583). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.919

Liu, X., Gao, Y., Zong, L., & Xu, B. (2024). Improve meta-learning for few-shot text classification with all you can acquire from the tasks. In Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 223–235). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.12

Lu, Y. J., & Li, C. T. (2020). GCAN: Graph-aware co-attention networks for explainable fake news detection on social media. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 505–514). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.48

Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (Vol. 30, pp. 4765–4774). Curran Associates, Inc.

Pekandi, L. A., Widjaja, R. G., Ananta, A., Harefa, J., & Jingga, K. (2025). Evaluating IndoBERT for Indonesian hoax news detection: A comparative study with ensemble and CNN-LSTM models. Procedia Computer Science, 269, 1625–1633. https://doi.org/10.1016/j.procs.2025.09.105

Prisscilya, V., & Girsang, A. S. (2024). Classification of Indonesia false news detection using BERTopic and IndoBERT. Jurnal Indonesia Sosial Teknologi, 5(8), 3913–3931. https://doi.org/10.59141/jist.v5i8.1310

Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). Why should I trust you? Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1135–1144). Association for Computing Machinery. https://doi.org/10.1145/2939672.2939778

Rogers, A., Kovaleva, O., & Rumshisky, A. (2020). A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8, 842–866. https://doi.org/10.1162/tacl_a_00349

Sehanobish, A., Kannan, K., Abraham, N., Das, A., & Odry, B. (2022). Meta-learning pathologies from radiology reports using variance aware prototypical networks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track (pp. 332–347). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-industry.34

Shu, K., Sliva, A., Wang, S., Tang, J., & Liu, H. (2017). Fake news detection on social media: A data mining perspective. ACM SIGKDD Explorations Newsletter, 19(1), 22–36. https://doi.org/10.1145/3137597.3137600

Simanjuntak, A., Lumbantoruan, R., Sianipar, K., Gultom, R., Simaremare, M., Situmeang, S., & Panggabean, E. (2024). Research and analysis of IndoBERT hyperparameter tuning in fake news detection. Jurnal Nasional Teknik Elektro dan Teknologi Informasi, 13(1), 60–67. https://doi.org/10.22146/jnteti.v13i1.8532

Snell, J., Swersky, K., & Zemel, R. S. (2017). Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems (Vol. 30, pp. 4077–4087). Curran Associates, Inc.

Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (Vol. 30, pp. 5998–6008). Curran Associates, Inc.

Wang, Y., Yao, Q., Kwok, J. T., & Ni, L. M. (2020). Generalizing from a few examples: A survey on few-shot learning. ACM Computing Surveys, 53(3), 1–34. https://doi.org/10.1145/3386252

Wen, X., Tan, W., & Weber, R. (2025). GAProtoNet: A multi-head graph attention-based prototypical network for interpretable text classification. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 9891–9901). Association for Computational Linguistics.

Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., Soleman, S., Mahendra, R., Fung, P., Bahar, S., & Purwarianti, A. (2020). IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing (pp. 843–857). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.aacl-main.85

Yefferson, D. Y., Lawijaya, V., & Girsang, A. S. (2024). Hybrid model: IndoBERT and long short-term memory for detecting Indonesian hoax news. IAES International Journal of Artificial Intelligence, 13(2), 1913–1924. https://doi.org/10.11591/ijai.v13.i2.pp1913-1924

Zhou, X., & Zafarani, R. (2020). A survey of fake news: Fundamental theories, detection methods, and opportunities. ACM Computing Surveys, 53(5), 1–40. https://doi.org/10.1145/3395046

Downloads

Published

2026-12-01

How to Cite

Christian, Y., Wilson, W., & Yulianto, A. (2026). Developing an Indonesian Fake News Detection System Using IndoBERT and ProtoNet for Few-Shot Learning. International Journal Software Engineering and Computer Science (IJSECS), 6(3), 1168-1184. https://doi.org/10.35870/ijsecs.v6i3.8156