How to process and extract text from image

Question

I'm trying to extract text from image using python cv2. The result is pathetic and I can't figure out a way to improve my code. I believe the image needs to be processed before the extraction of text but not sure how.

Sample image

I've tried to convert it into black and white but no luck.

import cv2
import os
import pytesseract
from PIL import Image
import time

pytesseract.pytesseract.tesseract_cmd='C:\\Program Files\\Tesseract-OCR\\tesseract.exe'

cam = cv2.VideoCapture(1,cv2.CAP_DSHOW)

cam.set(cv2.CAP_PROP_FRAME_WIDTH, 8000)
cam.set(cv2.CAP_PROP_FRAME_HEIGHT, 6000)

while True:
    return_value,image = cam.read()
    image=cv2.cvtColor(image,cv2.COLOR_BGR2GRAY)
    image = image[127:219, 508:722]
    #(thresh, image) = cv2.threshold(image, 128, 255, cv2.THRESH_BINARY | cv2.THRESH_OTSU)
    cv2.imwrite('test.jpg',image)
    print('Text detected: {}'.format(pytesseract.image_to_string(Image.open('test.jpg'))))
    time.sleep(2)

cam.release()
#os.system('del test.jpg')

You could also try to use EasyOCR. In my case the results were much better compared to Tesseract where it was just random text. At the moment, a custom model using EasyOCR cannot be trained. — Rishik Mani
– Rishik Mani, Commented Mar 31, 2021 at 14:07

nathancy · Accepted Answer · 2019-08-28 21:54:34Z

7

Preprocessing to clean the image before performing text extraction can help. Here's a simple approach

Convert image to grayscale and sharpen image
Adaptive threshold
Perform morpholgical operations to clean image
Invert image

First we convert to grayscale then sharpen the image using a sharpening kernel

Next we adaptive threshold to obtain a binary image

Now we perform morphological transformations to smooth the image

Finally we invert the image

import cv2
import numpy as np

image = cv2.imread('1.jpg')
gray = cv2.cvtColor(image, cv2.COLOR_BGR2GRAY)
sharpen_kernel = np.array([[-1,-1,-1], [-1,9,-1], [-1,-1,-1]])
sharpen = cv2.filter2D(gray, -1, sharpen_kernel)
thresh = cv2.threshold(sharpen, 0, 255, cv2.THRESH_BINARY_INV + cv2.THRESH_OTSU)[1]

kernel = cv2.getStructuringElement(cv2.MORPH_RECT, (3,3))
close = cv2.morphologyEx(thresh, cv2.MORPH_CLOSE, kernel, iterations=1)
result = 255 - close

cv2.imshow('sharpen', sharpen)
cv2.imshow('thresh', thresh)
cv2.imshow('close', close)
cv2.imshow('result', result)
cv2.waitKey()

answered Aug 28, 2019 at 21:54

nathancy

47k15 gold badges137 silver badges154 bronze badges

Sign up to request clarification or add additional context in comments.

3 Comments

idar Over a year ago

This actually helped me learn few good tricks about processing the image but still pytesseract is unable to extract any anything from the image.

nathancy Over a year ago

You may have to "cut up" the image and feed each line into pytesseract to get good results. Slicing the image horizontally in half to separate each word could help. Good luck!

idar Over a year ago

Thank you for suggesting this. I’ll try it and update the post.

Collectives™ on Stack Overflow

How to process and extract text from image

1 Answer 1

3 Comments

Your Answer

Linked

Hot Network Questions

Collectives™ on Stack Overflow

1 Answer 1

3 Comments

Your Answer

Sign up or log in

Post as a guest

Linked

Related