Directional Robustness Asymmetry in Indonesian Sarcasm Detection: A Cross-Platform Evaluation of IndoRoBERTa on X and Reddit

Authors

  • Marcellinus Brendan Hananta Universitas Kristen Satya Wacana
  • Wiwin Sulistyo Universitas Kristen Satya Wacana

DOI:

https://doi.org/10.35870/jtik.v11i1.7711

Keywords:

Cross-Domain Robustness, IdSarcasm, IndoRoBERTa, Natural Language Processing, Sarcasm Detection

Abstract

Sarcasm can reverse sentence polarity, making it one of the hardest problems in natural language processing and a persistent obstacle to reliable sentiment analysis on social media, where models are often deployed on platforms they were not trained on. Research on Indonesian sarcasm has largely stayed within a single domain or tested transfer in only one direction, leaving the cross-domain robustness of the IndoRoBERTa family on the IdSarcasm benchmark unclear. This study measures and compares the cross-domain robustness of IndoRoBERTa-small and IndoRoBERTa-base, with IndoBERT as a baseline, across X (Twitter) and Reddit, using the F1 gap (ΔF1) after cleaning anonymization artifacts, adding a data-size control, and running Stratified 5-Fold cross-validation. The results reveal a one-sided asymmetry: X → Reddit transfer is far more fragile (ΔF1 0.345–0.400) than Reddit → X (0.135–0.240), failing by missing sarcasm (false negatives) rather than over-flagging it (false positives). Practically, training on the context-rich domain (Reddit) and choosing IndoRoBERTa-base give the most reliable cross-platform sarcasm detection, with direction-specific mitigation recommended for robust deployment.

Downloads

Download data is not yet available.

Author Biographies

  • Marcellinus Brendan Hananta, Universitas Kristen Satya Wacana

    Program Studi Teknik Informatika, Fakultas Teknologi Informasi, Universitas Kristen Satya Wacana, Kota Salatiga, Jawa Tengah, Indonesia.

  • Wiwin Sulistyo, Universitas Kristen Satya Wacana

    Program Studi Teknik Informatika, Fakultas Teknologi Informasi, Universitas Kristen Satya Wacana, Kota Salatiga, Jawa Tengah, Indonesia.

References

Abdulkadirov, R., Lyakhov, P., & Nagornov, N. (2023). Survey of optimization algorithms in modern neural networks. Mathematics, 11(11), 2466. https://doi.org/10.3390/math11112466.

Abu Farha, I., Oprea, S. V., Wilson, S., & Magdy, W. (2022). SemEval-2022 Task 6: iSarcasmEval, intended sarcasm detection in English and Arabic. Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022), 802–814. https://doi.org/10.18653/v1/2022.semeval-1.111.

Ahmadian, H., Abidin, T. F., Riza, H., & Muchtar, K. (2024). Hybrid models for emotion classification and sentiment analysis in Indonesian language. Applied Computational Intelligence and Soft Computing, 2024(1), 2826773. https://doi.org/10.1155/2024/2826773.

Alfan Rosid, M., Oranova Siahaan, D., & Saikhu, A. (2024). Sarcasm detection in Indonesian-English code-mixed text using multihead attention-based convolutional and bi-directional GRU. IEEE Access, 12, 137063–137079. https://doi.org/10.1109/ACCESS.2024.3436107.

Ameur, A., Hamdi, S., & Yahia, S. B. (2023). Domain adaptation approach for Arabic sarcasm detection in hotel reviews based on hybrid learning. Procedia Computer Science, 225, 3898–3908. https://doi.org/10.1016/j.procs.2023.10.385.

Bouazizi, M., & Ohtsuki, T. (2022). Sarcasm over time and across platforms: Does the way we express sarcasm change? IEEE Access, 10, 55958–55987. https://doi.org/10.1109/ACCESS.2022.3174862.

Calderon, N., Porat, N., Ben-David, E., Chapanin, A., Gekhman, Z., Oved, N., Shalumov, V., & Reichart, R. (2024). Measuring the robustness of NLP models to domain shifts. arXiv. https://doi.org/10.48550/arXiv.2306.00168.

Chen, W., Lin, F., Li, G., & Liu, B. (2024). A survey of automatic sarcasm detection: Fundamental theories, formulation, datasets, detection methods, and opportunities. Neurocomputing, 578, 127428. https://doi.org/10.1016/j.neucom.2024.127428.

Chen, W., Lin, F., Zhang, X., Li, G., & Liu, B. (2022). Jointly learning sentimental clues and context incongruity for sarcasm detection. IEEE Access, 10, 48292–48300. https://doi.org/10.1109/ACCESS.2022.3169864.

Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186. https://doi.org/10.18653/v1/N19-1423.

Fitrianto, R. A., & Editya, A. S. (2024). Klasifikasi tweet sarkasme pada platform X menggunakan bidirectional encoder representations from transformers. Jurnal Teknologi Dan Sistem Informasi Bisnis, 6(3), 366–371. https://doi.org/10.47233/jteksis.v6i3.1344

Helal, N. A., Hassan, A., Badr, N. L., & Afify, Y. M. (2024). A contextual-based approach for sarcasm detection. Scientific Reports, 14(1), 15415. https://doi.org/10.1038/s41598-024-65217-8.

Khan, S., Qasim, I., Khan, W., Khan, A., Ali Khan, J., Qahmash, A., & Ghadi, Y. Y. (2024). An automated approach to identify sarcasm in low-resource language. PLOS ONE, 19(12), e0307186. https://doi.org/10.1371/journal.pone.0307186.

Khoee, A. G., Yu, Y., & Feldt, R. (2024). Domain generalization through meta-learning: A survey. Artificial Intelligence Review, 57(10), 285. https://doi.org/10.1007/s10462-024-10922-z.

Khotijah, S., Tirtawangsa, J., & Suryani, A. A. (2020). Using LSTM for context-based approach of sarcasm detection in Twitter. In Proceedings of the 11th International Conference on Advances in Information Technology (IAIT 2020). https://doi.org/10.1145/3406601.3406624

Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv. https://doi.org/10.48550/arXiv.1907.11692.

Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. arXiv. https://doi.org/10.48550/arXiv.1711.05101.

Mulia, B., Emily, M., Damario, V. A., & Hasani, M. F. (2025). Context-aware loss for Indonesian sarcasm detection: A comparative study with conventional imbalance-oriented objectives. 2025 5th International Conference on Intelligent Cybernetics Technology & Applications (ICICyTA), 882–888. https://doi.org/10.1109/ICICyTA68677.2025.11362728.

Oprea, S.-V., & Bâra, A. (2025). LLM-as-a-judge for sarcasm detection using supervised fine-tuning of transformers. Journal of King Saud University Computer and Information Sciences, 37(10), 357. https://doi.org/10.1007/s44443-025-00379-7.

Rahman, F., & Girsang, A. S. (2024). IndoBERTweet for sarcasm: Evaluating domain-adapted transformers for Indonesian Twitter sarcasm classification. Journal of Logistics, Informatics and Service Science, 11(2). https://doi.org/10.33168/JLISS.2024.0210.

Rainio, O., Teuho, J., & Klén, R. (2024). Evaluation metrics and statistical tests for machine learning. Scientific Reports, 14(1), 6086. https://doi.org/10.1038/s41598-024-56706-x.

Ranti, K. S., & Girsang, A. S. (2020). Indonesian sarcasm detection using convolutional neural network. International Journal of Emerging Trends in Engineering Research, 8(9), 4952–4955. https://doi.org/10.30534/ijeter/2020/10892020.

Sabrina, A. L., Wijaya, D. C. C., & Meiliana. (2026). From Twitter to Reddit: Cross-domain Indonesian sarcasm detection with pretrained transformers. 2026 International Conference on Current Research in Artificial Intelligence and Data Science (ICCRAIDS), 1–5. https://doi.org/10.1109/ICCRAIDS67816.2026.11519738.

Šandor, D., & Bagić Babac, M. (2024). Sarcasm detection in online comments using machine learning. Information Discovery and Delivery, 52(2), 213–226. https://doi.org/10.1108/IDD-01-2023-0002.

Suhartono, D., Wongso, W., & Tri Handoyo, A. (2024). IdSarcasm: Benchmarking and evaluating language models for Indonesian sarcasm detection. IEEE Access, 12, 87323–87332. https://doi.org/10.1109/ACCESS.2024.3416955.

Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., Soleman, S., Mahendra, R., Fung, P., Bahar, S., & Purwarianti, A. (2020). IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. arXiv. https://doi.org/10.48550/arXiv.2009.05387.

Downloads

Published

2027-01-01

Issue

Section

Computer & Communication Science

How to Cite

Hananta, M. B., & Sulistyo, W. (2027). Directional Robustness Asymmetry in Indonesian Sarcasm Detection: A Cross-Platform Evaluation of IndoRoBERTa on X and Reddit. Jurnal JTIK (Jurnal Teknologi Informasi Dan Komunikasi), 11(1), 172-183. https://doi.org/10.35870/jtik.v11i1.7711

Most read articles by the same author(s)