Sistem Pengenalan Citra Dokumen Teks Terdistorsi menjadi Teks Menggunakan Metode Deep Learning

Authors

  • Talenta Teholi Zalukhu Universitas Kristen Immanuel
  • Agustinus Rudatyo Himamunanto Universitas Kristen Immanuel
  • Haeni Budiati Universitas Kristen Immanuel

DOI:

https://doi.org/10.35870/jtik.v10i1.4700

Keywords:

OCR, DnCNN, Tesseract, Deblurring, Blurred Image, Deep Learning, Image Enhancement

Abstract

A common issue in document image processing is the inability of OCR systems to accurately read text from blurred images. This study aims to develop a deep learning-based OCR pipeline capable of recognizing text in blurred document images. The process begins with image enhancement using the DnCNN model for deblurring, followed by character segmentation and classification of A–Z characters using a CNN trained on the EMNIST Letters dataset. The recognized characters are then reconstructed into complete text. Experiments were conducted on 300 blurred images with varying levels of blur (low, medium, and high). Evaluation using PSNR and SSIM metrics showed improvements in image quality, with an average PSNR of 29,56 dB and SSIM of 0.89. Furthermore, the character classification accuracy reached 95.64%. Compared to the baseline (direct Tesseract OCR without deblurring), the proposed system showed a significant improvement in text readability. These results demonstrate the effectiveness of CNN-based approaches in enhancing OCR performance on blurred document images.

Downloads

Download data is not yet available.

Author Biographies

  • Talenta Teholi Zalukhu, Universitas Kristen Immanuel

    Program Studi Informatika, Fakultas Sains & Komputer, Universitas Kristen Immanuel, Kabupaten Sleman, Daerah Istimewa Yogyakarta, Indonesia.

  • Agustinus Rudatyo Himamunanto, Universitas Kristen Immanuel

    Program Studi Informatika, Fakultas Sains & Komputer, Universitas Kristen Immanuel, Kabupaten Sleman, Daerah Istimewa Yogyakarta, Indonesia.

  • Haeni Budiati, Universitas Kristen Immanuel

    Program Studi Informatika, Fakultas Sains & Komputer, Universitas Kristen Immanuel, Kabupaten Sleman, Daerah Istimewa Yogyakarta, Indonesia.

References

Baldominos, A., Saez, Y., & Isasi, P. (2019). A survey of handwritten character recognition with mnist and emnist. Applied sciences, 9(15), 3169.

Fadjeri, A., Asroriyah, A. M., & Rahmawati, A. (2022). Analisis Teks Bahasa Indonesia Dan Inggris Dari Sebuah Citra Menggunakan Pengolahan Citra Digital. Jurnal Teknologi Informasi Dan Komunikasi (TIKomSiN), 10(2), 42-46.

Giamiko, E. S., & Tjiong, E. (2024). Pengembangan Aplikasi Pengenalan Tulisan Tangan Abjad dan Angka Berbasis Convolutional Neural Network. KALBISCIENTIA Jurnal Sains dan Teknologi, 11(02), 22-30. https://doi.org/10.53008/kalbiscientia.v11i02.3626.

Hengaju, U., & Bal, B. K. (2023). Improving the Recognition Accuracy of Tesseract-OCR Engine on Nepali Text Images via Preprocessing. Advancement in Image Processing and Pattern Recognition, 3(2), 3.

Manurung, I. D. P., Hidayatno, A., & Setiyono, B. (2011). Pengenalan Teks Cetak Pada Citra Teks Biner. Universitas Diponegoro, Semarang.

Mohsenzadegan, K., Tavakkoli, V., & Kyamakya, K. (2022). Deep neural network concept for a blind enhancement of document-images in the presence of multiple distortions. Applied Sciences, 12(19), 9601.

Nugroho, A. W. (2024). PENGGUNAAN MACHINE LEARNING DALAM PENGENALAN TEKS DARI GAMBAR. Jurnal Dunia Data, 1(4).

Nugroho, B., & Puspaningrum, E. Y. (2021). Kinerja Metode CNN untuk Klasifikasi Pneumonia dengan Variasi Ukuran Citra Input. Jurnal Teknologi Informasi dan Ilmu Komputer (JTIIK), 8(3), 533-538.

Peryanto, A., Yudhana, A., & Umar, R. (2020). Rancang bangun klasifikasi citra dengan teknologi deep learning berbasis metode convolutional neural network. Format J. Ilm. Tek. Inform, 8(2), 138.

Pratiwi, A., Lestari, Y. D., & Lubis, Y. F. A. (2021, October). Analisis Kombinasi Vertical Projection Profile (VPP) Dan Top Down Profile (TDP) Dalam Segmentasi Karakter Pada Aplikasi OCR. In SEMINAR NASIONAL TEKNOLOGI INFORMASI & KOMUNIKASI (Vol. 1, No. 1, pp. 470-478).

Rizqi, A., & Aziz, I. S. (2019). Rancang Bangun Aplikasi Penerjemah Bahasa Jepang–Indonesia Menggunakan OCR Berbasis Android. SinarFe7, 2(1), 276-280.

Septiarini, A. (2016). Pengenalan Pola Pada Citra Digital Dengan Fitur Momen Invariant. Informatika Mulawarman: Jurnal Ilmiah Ilmu Komputer, 7(1), 8-11.

Shaliniswetha, S., & Mahaboob, S. T. (2022). RESIDUAL LEARNING BASED IMAGE DENOISING AND COMPRESSION USING DNCNN. ICTACT Journal on Image & Video Processing, 13(2).

Siliwangi, A. K., & Prabowo, Y. D. (2022). Pencarian Informasi Berbasis Teks dalam Komik Digital Menggunakan OCR. KALBISIANA Jurnal Sains, Bisnis dan Teknologi, 8(2), 1886-1894.

Sulistiyo, M. F. (2022). Penerjemah Bahasa Inggris-Indonesia Berbasis Mobile Menggunakan Optical Character Recognition & Text To Speech (Doctoral dissertation, UBP Karawang).

Wijaya, I., & Lubis, C. (2022). Pengimplementasian Ocr Menggunakan Cnn Untuk Ekstraksi Teks Pada Gambar. Jurnal Ilmu Komputer dan Sistem Informasi, 10(1). https://doi.org/10.24912/jiksi.v10i1.17836.

Downloads

Published

2026-01-01

Issue

Section

Computer & Communication Science

How to Cite

Zalukhu, T. T., Himamunanto, A. R., & Budiati, H. (2026). Sistem Pengenalan Citra Dokumen Teks Terdistorsi menjadi Teks Menggunakan Metode Deep Learning. Jurnal JTIK (Jurnal Teknologi Informasi Dan Komunikasi), 10(1), 1-14. https://doi.org/10.35870/jtik.v10i1.4700

Similar Articles

You may also start an advanced similarity search for this article.