Directional Robustness Asymmetry in Indonesian Sarcasm Detection: A Cross-Platform Evaluation of IndoRoBERTa on X and Reddit
DOI:
https://doi.org/10.35870/jtik.v11i1.7711Keywords:
Cross-Domain Robustness, IdSarcasm, IndoRoBERTa, Natural Language Processing, Sarcasm DetectionAbstract
Sarcasm can reverse sentence polarity, making it one of the hardest problems in natural language processing and a persistent obstacle to reliable sentiment analysis on social media, where models are often deployed on platforms they were not trained on. Research on Indonesian sarcasm has largely stayed within a single domain or tested transfer in only one direction, leaving the cross-domain robustness of the IndoRoBERTa family on the IdSarcasm benchmark unclear. This study measures and compares the cross-domain robustness of IndoRoBERTa-small and IndoRoBERTa-base, with IndoBERT as a baseline, across X (Twitter) and Reddit, using the F1 gap (ΔF1) after cleaning anonymization artifacts, adding a data-size control, and running Stratified 5-Fold cross-validation. The results reveal a one-sided asymmetry: X → Reddit transfer is far more fragile (ΔF1 0.345–0.400) than Reddit → X (0.135–0.240), failing by missing sarcasm (false negatives) rather than over-flagging it (false positives). Practically, training on the context-rich domain (Reddit) and choosing IndoRoBERTa-base give the most reliable cross-platform sarcasm detection, with direction-specific mitigation recommended for robust deployment.
Downloads
References
Abdulkadirov, R., Lyakhov, P., & Nagornov, N. (2023). Survey of optimization algorithms in modern neural networks. Mathematics, 11(11), 2466. https://doi.org/10.3390/math11112466.
Abu Farha, I., Oprea, S. V., Wilson, S., & Magdy, W. (2022). SemEval-2022 Task 6: iSarcasmEval, intended sarcasm detection in English and Arabic. Proceedings of the 16th International Workshop on Semantic Evaluation (SemEval-2022), 802–814. https://doi.org/10.18653/v1/2022.semeval-1.111.
Ahmadian, H., Abidin, T. F., Riza, H., & Muchtar, K. (2024). Hybrid models for emotion classification and sentiment analysis in Indonesian language. Applied Computational Intelligence and Soft Computing, 2024(1), 2826773. https://doi.org/10.1155/2024/2826773.
Alfan Rosid, M., Oranova Siahaan, D., & Saikhu, A. (2024). Sarcasm detection in Indonesian-English code-mixed text using multihead attention-based convolutional and bi-directional GRU. IEEE Access, 12, 137063–137079. https://doi.org/10.1109/ACCESS.2024.3436107.
Ameur, A., Hamdi, S., & Yahia, S. B. (2023). Domain adaptation approach for Arabic sarcasm detection in hotel reviews based on hybrid learning. Procedia Computer Science, 225, 3898–3908. https://doi.org/10.1016/j.procs.2023.10.385.
Bouazizi, M., & Ohtsuki, T. (2022). Sarcasm over time and across platforms: Does the way we express sarcasm change? IEEE Access, 10, 55958–55987. https://doi.org/10.1109/ACCESS.2022.3174862.
Calderon, N., Porat, N., Ben-David, E., Chapanin, A., Gekhman, Z., Oved, N., Shalumov, V., & Reichart, R. (2024). Measuring the robustness of NLP models to domain shifts. arXiv. https://doi.org/10.48550/arXiv.2306.00168.
Chen, W., Lin, F., Li, G., & Liu, B. (2024). A survey of automatic sarcasm detection: Fundamental theories, formulation, datasets, detection methods, and opportunities. Neurocomputing, 578, 127428. https://doi.org/10.1016/j.neucom.2024.127428.
Chen, W., Lin, F., Zhang, X., Li, G., & Liu, B. (2022). Jointly learning sentimental clues and context incongruity for sarcasm detection. IEEE Access, 10, 48292–48300. https://doi.org/10.1109/ACCESS.2022.3169864.
Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of deep bidirectional transformers for language understanding. Proceedings of the 2019 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long and Short Papers), 4171–4186. https://doi.org/10.18653/v1/N19-1423.
Fitrianto, R. A., & Editya, A. S. (2024). Klasifikasi tweet sarkasme pada platform X menggunakan bidirectional encoder representations from transformers. Jurnal Teknologi Dan Sistem Informasi Bisnis, 6(3), 366–371. https://doi.org/10.47233/jteksis.v6i3.1344
Helal, N. A., Hassan, A., Badr, N. L., & Afify, Y. M. (2024). A contextual-based approach for sarcasm detection. Scientific Reports, 14(1), 15415. https://doi.org/10.1038/s41598-024-65217-8.
Khan, S., Qasim, I., Khan, W., Khan, A., Ali Khan, J., Qahmash, A., & Ghadi, Y. Y. (2024). An automated approach to identify sarcasm in low-resource language. PLOS ONE, 19(12), e0307186. https://doi.org/10.1371/journal.pone.0307186.
Khoee, A. G., Yu, Y., & Feldt, R. (2024). Domain generalization through meta-learning: A survey. Artificial Intelligence Review, 57(10), 285. https://doi.org/10.1007/s10462-024-10922-z.
Khotijah, S., Tirtawangsa, J., & Suryani, A. A. (2020). Using LSTM for context-based approach of sarcasm detection in Twitter. In Proceedings of the 11th International Conference on Advances in Information Technology (IAIT 2020). https://doi.org/10.1145/3406601.3406624
Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., Levy, O., Lewis, M., Zettlemoyer, L., & Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach. arXiv. https://doi.org/10.48550/arXiv.1907.11692.
Loshchilov, I., & Hutter, F. (2019). Decoupled weight decay regularization. arXiv. https://doi.org/10.48550/arXiv.1711.05101.
Mulia, B., Emily, M., Damario, V. A., & Hasani, M. F. (2025). Context-aware loss for Indonesian sarcasm detection: A comparative study with conventional imbalance-oriented objectives. 2025 5th International Conference on Intelligent Cybernetics Technology & Applications (ICICyTA), 882–888. https://doi.org/10.1109/ICICyTA68677.2025.11362728.
Oprea, S.-V., & Bâra, A. (2025). LLM-as-a-judge for sarcasm detection using supervised fine-tuning of transformers. Journal of King Saud University Computer and Information Sciences, 37(10), 357. https://doi.org/10.1007/s44443-025-00379-7.
Rahman, F., & Girsang, A. S. (2024). IndoBERTweet for sarcasm: Evaluating domain-adapted transformers for Indonesian Twitter sarcasm classification. Journal of Logistics, Informatics and Service Science, 11(2). https://doi.org/10.33168/JLISS.2024.0210.
Rainio, O., Teuho, J., & Klén, R. (2024). Evaluation metrics and statistical tests for machine learning. Scientific Reports, 14(1), 6086. https://doi.org/10.1038/s41598-024-56706-x.
Ranti, K. S., & Girsang, A. S. (2020). Indonesian sarcasm detection using convolutional neural network. International Journal of Emerging Trends in Engineering Research, 8(9), 4952–4955. https://doi.org/10.30534/ijeter/2020/10892020.
Sabrina, A. L., Wijaya, D. C. C., & Meiliana. (2026). From Twitter to Reddit: Cross-domain Indonesian sarcasm detection with pretrained transformers. 2026 International Conference on Current Research in Artificial Intelligence and Data Science (ICCRAIDS), 1–5. https://doi.org/10.1109/ICCRAIDS67816.2026.11519738.
Šandor, D., & Bagić Babac, M. (2024). Sarcasm detection in online comments using machine learning. Information Discovery and Delivery, 52(2), 213–226. https://doi.org/10.1108/IDD-01-2023-0002.
Suhartono, D., Wongso, W., & Tri Handoyo, A. (2024). IdSarcasm: Benchmarking and evaluating language models for Indonesian sarcasm detection. IEEE Access, 12, 87323–87332. https://doi.org/10.1109/ACCESS.2024.3416955.
Wilie, B., Vincentio, K., Winata, G. I., Cahyawijaya, S., Li, X., Lim, Z. Y., Soleman, S., Mahendra, R., Fung, P., Bahar, S., & Purwarianti, A. (2020). IndoNLU: Benchmark and resources for evaluating Indonesian natural language understanding. arXiv. https://doi.org/10.48550/arXiv.2009.05387.
Downloads
Published
Issue
Section
License
Copyright (c) 2027 Marcellinus Brendan Hananta, Wiwin Sulistyo

This work is licensed under a Creative Commons Attribution 4.0 International License.
Authors who publish with this journal agree to the following terms:
1. Copyright Retention and Open Access License
Authors retain copyright of their work and grant the journal non-exclusive right of first publication under the Creative Commons Attribution 4.0 International License (CC BY 4.0).
This license allows unrestricted use, distribution, and reproduction in any medium, provided the original work is properly cited.
2. Rights Granted Under CC BY 4.0
Under this license, readers are free to:
- Share — copy and redistribute the material in any medium or format
- Adapt — remix, transform, and build upon the material for any purpose, including commercial use
- No additional restrictions — the licensor cannot revoke these freedoms as long as license terms are followed
3. Attribution Requirements
All uses must include:
- Proper citation of the original work
- Link to the Creative Commons license
- Indication if changes were made to the original work
- No suggestion that the licensor endorses the user or their use
4. Additional Distribution Rights
Authors may:
- Deposit the published version in institutional repositories
- Share through academic social networks
- Include in books, monographs, or other publications
- Post on personal or institutional websites
Requirement: All additional distributions must maintain the CC BY 4.0 license and proper attribution.
5. Self-Archiving and Pre-Print Sharing
Authors are encouraged to:
- Share pre-prints and post-prints online
- Deposit in subject-specific repositories (e.g., arXiv, bioRxiv)
- Engage in scholarly communication throughout the publication process
6. Open Access Commitment
This journal provides immediate open access to all content, supporting the global exchange of knowledge without financial, legal, or technical barriers.
