Extracting the text from a PDF document is surprisingly difficult. If you haven’t worked with PDF files before, you’ll be surprised to learn that a PDF document does not contain any text at all. Additionally, there are two main types of PDF files: digital PDF files and scanned PDF files.
Extracting the text from a digital PDF file is tricky but there are dozens of techniques and tools that can accurately extract the text. But accurately extracting the text from a scanned PDF file (typically from a fax machine or converted from a photo of some sort) is nearly impossible using standard tools. The extracted text will almost always have many misspellings and errors.

This is the demo source scanned PDF. Extracting text from a scanned PDF is extremely difficult.
I was working on a project with thousands of PDF files where I needed to accurately extract the text. Some of the PDF files are digital and some are scanned. After many failures, I finally came up with a working solution. My solution has two phases. First, I programmatically extract the text from the source PDF document using the PyMuPDF library which has built-in OCR functionality. This gives rough text that is pretty good but there are quite a few spelling errors, missing spaces and so on. Second, I programmatically send the rough text to the OpenAI (aka ChatGPT) system and instruct the system to correct the rough text.
The two-phase extraction system works for scanned or digital PDF files.
For my demo, I downloaded a scanned PDF document I found on the Internet at nlsblog.org/wp-content/uploads/2020/06/image-based-pdf-sample.pdf.
To use the PyMuPDF library for the first phase of text extraction, I had to download and install the industry-standard Tesseract OCR engine. The Windows executable installer is tesseract-ocr-w64-setup-5.5.0.20241111.exe at github.com/UB-Mannheim/tesseract/wiki. The installer did not automatically set my environment PATH variable, so I had to do that manually (C:\Program Files\Tesseract-OCR).
After the Tesseract engine was installed, I installed the PyMuPDF library by issuing the command “pip install pymupdf”. For reasons that are completely unknown to me, PyMuPDF is also called “fitz”.
To use the OpenAI API, I had to create an OpenAI API account (not the same as a ChatGPT account). Using the OpenAI API is not free (it was at one time), but experiments only cost a few cents. After the API account was created, I generated a secret key, which is needed to use the API. Then I installed the openai library by issuing the command “pip install openai”.
Note: Instead of using PyMuPDF followed by OpenAI, I tried using OpenAI directly. Sometimes that simpler approach worked, but sometimes (usually) it failed on scanned PDF documents.
The demo program has the “riding a bicycle” characteristic — it was very difficult to create at first, but now it looks simple to me. So, anyway, if you’re reading this blog post as a resource, examine the code very carefully because it’s more subtle than it might first appear.
Instead of using the OpenAI API to correct spelling, an alternative approach for phase two is to use a dedicated spell checking such as the Python language pyspellchecker library. This alternative is more difficult but easier to customize for words like proprietary acronyms and people’s names. I’ll post a version of that when I get a chance.

A couple of entertaining old scanned newspaper headlines.
Demo program:
# extract_text_from_pdf.py
# extract text from a PDF file using PyMuPDF + OpenAI
# uses OCR so works with both scanned and digital PDFs
# requires tesseract OCR engine
# Go to:
# https://github.com/UB-Mannheim/tesseract/wiki
# Download and then double-click install
# tesseract-ocr-w64-setup-5.5.0.20241111.exe
# Add C:\Program Files\Tesseract-OCR to PATH
# pip install PyMuPDF
# pip install openai
import pymupdf # aka fitz
from openai import OpenAI
# -----------------------------------------------------------
def pdf_to_text(pdf_path, txt_path):
try:
# 1. get rough text from scanned PDF
doc = pymupdf.open(pdf)
txt = ""
for page in doc:
tp = page.get_textpage_ocr()
ocr_text = tp.extractTEXT() # aka get_text("text")
txt += ocr_text
doc.close()
# 2. send the rough text to openai API
key = "sk-proj-_AX7bGTXUwg-qojh2T5Z2CVXrox" + \
"xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx" + \
"xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx" + \
"xxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxxx"
client = OpenAI(api_key=key)
command = "The text might have has missing spaces " + \
"or incorrectly spelled words. Correct the " + \
"spelling. Don't correct any grammar. Respond " + \
"only with the corrected text."
response = client.responses.create(
model = "gpt-4.1",
input = [
{ "role": "system",
"content": "You correct text spelling." },
{ "role": "user",
"content": txt },
{ "role": "user",
"content": command },
],
temperature = 0.1, # not creative
top_p = 1.0, # default max explore
max_output_tokens = 10000,
)
# print(response.output_text)
# 3. write the corrected text to file
with open(txt_path, "w", encoding="utf-8") as tf:
tf.write(response.output_text)
except Exception as e:
print("Fatal error " + str(e))
# -----------------------------------------------------------
# the technique works with digital PDFs too
# pdf = "./PDFs/venus_pink.pdf" # digital PDF
# txt = "./TextFiles/venus_pink.txt"
pdf = "./PDFs/scanned_example.pdf" # scanned
txt = "./TextFiles/scanned_example.txt"
print("\nConverting " + str(pdf) + " to " + str(txt))
pdf_to_text(pdf, txt)
print("Done ")

.NET Test Automation Recipes
Software Testing
SciPy Programming Succinctly
Keras Succinctly
R Programming
Visual Studio Live
Microsoft MLADS Conference
DevIntersection Conference
Machine Learning Week
Ai4 Conference
G2E Conference
iSC West Conference
You must be logged in to post a comment.