Developing an Indonesian Fake News Detection System Using IndoBERT and ProtoNet for Few-Shot Learning
DOI:
https://doi.org/10.35870/ijsecs.v6i3.8156Keywords:
Fake News Detection, IndoBERT, Prototypical Network, Few-Shot Learning, Explainable Artificial IntelligenceAbstract
Fake news dissemination through digital media has become a critical concern because inaccurate information may spread swiftly and influence public comprehension, social behavior, and decision-making. This study developed an Indonesian fake news detection web application employing fine-tuned IndoBERT and an IndoBERT-based Prototypical Network (ProtoNet) under few-shot learning scenarios. The research followed the Cross Industry Standard Process for Data Mining framework, consisting of business understanding, data understanding, data preparation, modeling, evaluation, and deployment. The dataset was obtained from the publicly available Deteksi Berita Hoaks Indo Dataset on Kaggle and contained 23,944 Indonesian news records, consisting of 11,200 factual and 12,744 hoax articles. Three few-shot scenarios were evaluated: 50-shot, 100-shot, and 300-shot per class. IndoBERT was fine-tuned as a supervised classification baseline, while ProtoNet used IndoBERT embeddings to construct class prototypes and classify texts based on Euclidean distance. The best performance was achieved by fine-tuned IndoBERT in the 300-shot scenario, with 97.12% accuracy and a 97.10% macro F1-score. ProtoNet achieved 92.65% accuracy and a 92.63% macro F1-score in the same scenario. Even in the most constrained 50-shot setting, ProtoNet achieved an F1-score of 90.82%, closely trailing fine-tuned IndoBERT at 91.17% and demonstrating strong metric stability. The web application supports manual text input, URL-based article extraction, model selection, confidence scores, and word-level occlusion explanations. These results demonstrate the effectiveness of IndoBERT for Indonesian fake news detection and the feasibility of prototype-based classification under limited-data conditions. Overall, this work bridges the gap between deep contextual representation, data-efficient learning, and practical, interpretable deployment for public media verification.
Downloads
References
Arrieta, A. B., Díaz-Rodríguez, N., Del Ser, J., Bennetot, A., Tabik, S., Barbado, A., García, S., Gil-López, S., Molina, D., Benjamins, R., Chatila, R., & Herrera, F. (2020). Explainable artificial intelligence (XAI): Concepts, taxonomies, opportunities and challenges toward responsible AI. Information Fusion, 58, 82–115. https://doi.org/10.1016/j.inffus.2019.12.012
Azizah, S. F. N., Cahyono, H. D., Sihwi, S. W., & Widiarto, W. (2023). Performance analysis of transformer based models (BERT, ALBERT and RoBERTa) in fake news detection. In 2023 6th International Conference on Information and Communications Technology (ICOIACT) (pp. 425–430). IEEE. https://doi.org/10.1109/ICOIACT59844.2023.10455849
Bao, Y., Wu, M., Chang, S., & Barzilay, R. (2019). Few-shot text classification with distributional signatures. arXiv preprint arXiv:1908.06039. https://doi.org/10.48550/arXiv.1908.06039
Christian, Y., & Yap Rui Qi, K. O. (2022). Penerapan K-Means pada segmentasi pasar untuk riset pemasaran pada startup early stage dengan menggunakan CRISP-DM. JURIKOM (Jurnal Riset Komputer), 9(4), 966–973. https://doi.org/10.30865/jurikom.v9i4.4486
Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (pp. 4171–4186). Association for Computational Linguistics. https://doi.org/10.18653/v1/N19-1423
Dopierre, T., Gravier, C., & Logerais, W. (2021). A neural few-shot text classification reality check. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume (pp. 935–943). Association for Computational Linguistics. https://doi.org/10.18653/v1/2021.eacl-main.79
Erlina, E., Julyanto, J., Rustandi, J., Alexander, A., Francisco, L., Ma’muriyah, N., & Sabariman, S. (2023). Penerapan artificial intelligence pada aplikasi chatbot sebagai sistem pelayanan dan informasi online pada sekolah. Journal of Information System and Technology, 4(3), 610–621. https://doi.org/10.37253/joint.v4i3.6296
Fakhruzzaman, M. N., & Gunawan, S. W. (2021). Web-based application for detecting Indonesian clickbait headlines using IndoBERT. arXiv preprint arXiv:2102.10601. https://doi.org/10.48550/arXiv.2102.10601
Fathin, M. A., Sibaroni, Y., & Prasetyowati, S. S. (2024). Handling imbalance dataset on hoax Indonesian political news classification using IndoBERT and random sampling. Jurnal Media Informatika Budidarma, 8(1), 352–360. https://doi.org/10.30865/mib.v8i1.7099
Fawaid, J., Awalina, A., Krisnabayu, R. Y., & Yudistira, N. (2021). Indonesia’s fake news detection using transformer network. In Proceedings of the 6th International Conference on Sustainable Information Engineering and Technology (pp. 247–251). Association for Computing Machinery. https://doi.org/10.1145/3479645.3479666
Geng, R., Li, B., Li, Y., Zhu, X., Jian, P., & Sun, J. (2019). Induction networks for few-shot text classification. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Language Processing (pp. 3904–3913). Association for Computational Linguistics. https://doi.org/10.18653/v1/D19-1403
Han, L., Zhang, X., Zhou, Z., & Liu, Y. (2024). A multifaceted reasoning network for explainable fake news detection. Information Processing & Management, 61(6), Article 103822. https://doi.org/10.1016/j.ipm.2024.103822
Jocelynne, C., Wijayakusuma, I. G. N. L., & Harini, L. P. I. (2025). Detection of political hoax news using fine-tuning IndoBERT. Journal of Applied Informatics and Computing, 9(2), 354–360. https://doi.org/10.30871/jaic.v9i2.8989
Koto, F., Rahimi, A., Lau, J. H., & Baldwin, T. (2020). IndoLEM and IndoBERT: A benchmark dataset and pre-trained language model for Indonesian NLP. In Proceedings of the 28th International Conference on Computational Linguistics (pp. 757–770). International Committee on Computational Linguistics. https://doi.org/10.18653/v1/2020.coling-main.66
Lei, S., Zhang, X., He, J., Chen, F., & Lu, C.-T. (2023). TART: Improved few-shot text classification using task-adaptive reference transformation. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers) (pp. 11014–11026). Association for Computational Linguistics. https://doi.org/10.18653/v1/2023.acl-long.617
Liu, H., Wang, W., Li, H., & Li, H. (2024). TELLER: A trustworthy framework for explainable, generalizable and controllable fake news detection. In Findings of the Association for Computational Linguistics: ACL 2024 (pp. 15556–15583). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-acl.919
Liu, X., Gao, Y., Zong, L., & Xu, B. (2024). Improve meta-learning for few-shot text classification with all you can acquire from the tasks. In Findings of the Association for Computational Linguistics: EMNLP 2024 (pp. 223–235). Association for Computational Linguistics. https://doi.org/10.18653/v1/2024.findings-emnlp.12
Lu, Y. J., & Li, C. T. (2020). GCAN: Graph-aware co-attention networks for explainable fake news detection on social media. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics (pp. 505–514). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.acl-main.48
Lundberg, S. M., & Lee, S. I. (2017). A unified approach to interpreting model predictions. In Advances in Neural Information Processing Systems (Vol. 30, pp. 4765–4774). Curran Associates, Inc.
Pekandi, L. A., Widjaja, R. G., Ananta, A., Harefa, J., & Jingga, K. (2025). Evaluating IndoBERT for Indonesian hoax news detection: A comparative study with ensemble and CNN-LSTM models. Procedia Computer Science, 269, 1625–1633. https://doi.org/10.1016/j.procs.2025.09.105
Prisscilya, V., & Girsang, A. S. (2024). Classification of Indonesia false news detection using BERTopic and IndoBERT. Jurnal Indonesia Sosial Teknologi, 5(8), 3913–3931. https://doi.org/10.59141/jist.v5i8.1310
Ribeiro, M. T., Singh, S., & Guestrin, C. (2016). Why should I trust you? Explaining the predictions of any classifier. In Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining (pp. 1135–1144). Association for Computing Machinery. https://doi.org/10.1145/2939672.2939778
Rogers, A., Kovaleva, O., & Rumshisky, A. (2020). A primer in BERTology: What we know about how BERT works. Transactions of the Association for Computational Linguistics, 8, 842–866. https://doi.org/10.1162/tacl_a_00349
Sehanobish, A., Kannan, K., Abraham, N., Das, A., & Odry, B. (2022). Meta-learning pathologies from radiology reports using variance aware prototypical networks. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Industry Track (pp. 332–347). Association for Computational Linguistics. https://doi.org/10.18653/v1/2022.emnlp-industry.34
Shu, K., Sliva, A., Wang, S., Tang, J., & Liu, H. (2017). Fake news detection on social media: A data mining perspective. ACM SIGKDD Explorations Newsletter, 19(1), 22–36. https://doi.org/10.1145/3137597.3137600
Simanjuntak, A., Lumbantoruan, R., Sianipar, K., Gultom, R., Simaremare, M., Situmeang, S., & Panggabean, E. (2024). Research and analysis of IndoBERT hyperparameter tuning in fake news detection. Jurnal Nasional Teknik Elektro dan Teknologi Informasi, 13(1), 60–67. https://doi.org/10.22146/jnteti.v13i1.8532
Snell, J., Swersky, K., & Zemel, R. S. (2017). Prototypical networks for few-shot learning. In Advances in Neural Information Processing Systems (Vol. 30, pp. 4077–4087). Curran Associates, Inc.
Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, Ł., & Polosukhin, I. (2017). Attention is all you need. In Advances in Neural Information Processing Systems (Vol. 30, pp. 5998–6008). Curran Associates, Inc.
Wang, Y., Yao, Q., Kwok, J. T., & Ni, L. M. (2020). Generalizing from a few examples: A survey on few-shot learning. ACM Computing Surveys, 53(3), 1–34. https://doi.org/10.1145/3386252
Wen, X., Tan, W., & Weber, R. (2025). GAProtoNet: A multi-head graph attention-based prototypical network for interpretable text classification. In Proceedings of the 31st International Conference on Computational Linguistics (pp. 9891–9901). Association for Computational Linguistics.
Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., Soleman, S., Mahendra, R., Fung, P., Bahar, S., & Purwarianti, A. (2020). IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. In Proceedings of the 1st Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics and the 10th International Joint Conference on Natural Language Processing (pp. 843–857). Association for Computational Linguistics. https://doi.org/10.18653/v1/2020.aacl-main.85
Yefferson, D. Y., Lawijaya, V., & Girsang, A. S. (2024). Hybrid model: IndoBERT and long short-term memory for detecting Indonesian hoax news. IAES International Journal of Artificial Intelligence, 13(2), 1913–1924. https://doi.org/10.11591/ijai.v13.i2.pp1913-1924
Zhou, X., & Zafarani, R. (2020). A survey of fake news: Fundamental theories, detection methods, and opportunities. ACM Computing Surveys, 53(5), 1–40. https://doi.org/10.1145/3395046
Downloads
Published
Issue
Section
License
Copyright (c) 2026 Yefta Christian, Wilson Wilson, Andik Yulianto

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
1. Copyright Retention and Open Access License
Authors retain copyright of their work and grant the journal non-exclusive right of first publication under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
This license allows unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
2. Rights Granted Under CC BY 4.0
Under this license, readers are free to:
- Share — copy and redistribute the material in any medium or format
- Adapt — remix, transform, and build upon the material for any purpose, including commercial use
- No additional restrictions — the licensor cannot revoke these freedoms as long as license terms are followed
3. Attribution Requirements
All uses must include:
- Proper citation of the original work
- Link to the Creative Commons license
- Indication if changes were made to the original work
- No suggestion that the licensor endorses the user or their use
4. Additional Distribution Rights
Authors may:
- Deposit the published version in institutional repositories
- Share through academic social networks
- Include in books, monographs, or other publications
- Post on personal or institutional websites
Requirement: All additional distributions must maintain the CC BY 4.0 license and proper attribution.
5. Self-Archiving and Pre-Print Sharing
Authors are encouraged to:
- Share pre-prints and post-prints online
- Deposit in subject-specific repositories (e.g., arXiv, bioRxiv)
- Engage in scholarly communication throughout the publication process
6. Open Access Commitment
This journal provides immediate open access to all content, supporting the global exchange of knowledge without financial, legal, or technical barriers.
