PDF OCR

Choose a scanned PDF

PDF only · 20 MB · 40 pages

or drop one PDF here

Use this tool when a PDF contains scanned page images rather than selectable text. It renders the pages and runs optical character recognition (OCR) in your browser, then shows the recognised plain text for you to review, copy or save as a .txt file. The source PDF stays on your device. This tool does not create a searchable PDF or preserve the page layout, and OCR output can contain mistakes.

How to extract text from a scanned PDF

  1. 1

    Choose the document language

    Select the main language used in the printed text. The choice determines which OCR model the browser uses.

  2. 2

    Open the PDF

    Choose a PDF of up to 20 MB and 40 pages. The source file remains in your browser and is not sent to our server.

  3. 3

    Wait for recognition

    The browser renders each page and reads it with the selected OCR model. The model files may need to download before recognition starts.

  4. 4

    Check the result

    Compare names, numbers and important passages with the original pages. Treat OCR output as a draft and check it against the original.

  5. 5

    Copy or download the text

    Copy keeps the result in the browser. Download TXT locally saves a plain-text file without uploading it. Creating a protected download page is a separate, optional action.

What this tool produces

The output is plain text with a page marker before the recognised text from each page. It does not:

  • add a hidden text layer to the PDF
  • recreate columns, tables, forms or the original typography
  • preserve the visual page layout
  • translate the document
  • verify that the recognised text is correct

If you need a PDF that keeps the scanned pages and adds searchable text behind them, use a dedicated searchable-PDF or desktop OCR application.

Documents that work best

OCR is intended for printed text. A straight, clear scan with readable contrast is easier to recognise than a blurred photograph, faded fax or damaged page. Decorative type, handwriting, mathematical notation, closely packed columns and complex tables can produce poor or disordered text.

The tool accepts PDFs up to 20 MB and 40 pages. Recognition happens one page at a time and can use substantial memory, especially on large or high-resolution pages. If the browser runs short of memory, split the PDF into smaller parts and process them separately.

Supported OCR languages

You can choose English, Spanish, French, German, Italian, Portuguese, Dutch, Russian, Japanese, Chinese (Simplified), Korean or Arabic.

Choose the main language of the document before opening the file. A page that mixes languages or scripts may need separate passes, and the result can still confuse characters that look alike.

Review the recognised text

Always compare the result with the original PDF before relying on it. Pay particular attention to:

  • names, addresses and identification numbers
  • dates, amounts, decimal marks and minus signs
  • legal, medical, financial or safety instructions
  • page breaks and reading order in multi-column documents

Do not treat OCR output as an authoritative copy of a critical document.

File handling and downloads

The source PDF is read locally in your browser. It is not included in the funnel URL or sent to our server. Your browser may download the PDF rendering and OCR libraries and the selected language model over the network.

Copy keeps the recognised text in the browser and writes it to your clipboard. Download TXT locally builds a derived .txt file and saves it without uploading it. If you choose Create a download page, the tool uploads only that text file to create a protected results link, where it can be retained for up to 7 days. If that optional upload fails, the local download remains available.

Frequently Asked Questions

No. The source PDF remains in your browser while the pages are rendered and recognised. The browser may download the required PDF rendering and OCR libraries and the selected language model. Copy and Download TXT locally do not upload the recognised text. Only the separate Create a download page action uploads the derived plain-text file.

No. The result is plain text grouped by page. The tool does not modify the original PDF, add a hidden text layer or preserve its layout.

You should review it against the original pages. OCR can confuse characters, omit text and change reading order. Check every detail before using the output for legal, medical, financial, safety or other critical work.

It is designed for printed text. Handwriting, equations and decorative type can be unreliable. Characters inside a table may be recognised, but the row and column structure is not preserved in the plain-text result.

English, Spanish, French, German, Italian, Portuguese, Dutch, Russian, Japanese, Chinese (Simplified), Korean and Arabic. Select the main language before opening the PDF.

The PDF can be up to 20 MB and 40 pages. Large or detailed pages may still exceed the memory available to your browser, so splitting a difficult document into smaller files can help.

Download TXT locally creates a .txt file from the recognised text without uploading it. If you choose Create a download page, the derived file is uploaded to a protected results link that can be retained for up to 7 days. The local download remains available if that optional upload fails.

Related Tools