OCR

PDF OCR

Recognize scan pixels and rebuild searchable PDFs at original displayed sizes

Render & Extract
🔒 100% client-side — your data never leaves this page
Maintained by Evan•Updated: September 30, 2026
⏳Loading tool…

About this tool

Render selected visible PDF pages and recognize their pixels with a caller-owned Tesseract worker. Choose English, Simplified Chinese or both; automatic, single-block or sparse-text layout; and 1.5×, 2× or 2.5× rendering. Full TXT and JSON exports include source page numbers and OCR confidence estimates. The optional searchable PDF scales the raster image and hidden text together to the original physical displayed page dimensions, including CropBox, rotation and UserUnit. It does not preserve original vectors, forms, annotations, links, bookmarks or signatures. Input: 75 MiB and 1000 pages; select at most 40 pages, 16 million pixels each and 80 million total. Language files download on demand; documents stay local. All-empty or failed jobs create no completed/partial export.

Suggested Workflow

Scenario Recipes

01

Create searchable text from an image-only invoice

Goal: Compare recognition at two resolutions without changing physical page size.

  1. Load the two-page image-only example; select 1, choose English, Single text block and 1.5×, and enable searchable PDF.
  2. Run OCR, verify Invoice 2026-0930 and Total USD 128.50, then download TXT, JSON and PDF.
  3. Change to 2.5× and run again. Check the amount, the PDF text selection and the original 800 × 250 pt page size in both files.

Result: A reviewed searchable raster PDF and full text exports; higher scale changes raster resolution, not physical paper size.

Failure Clinic (Common Pitfalls)

The new PDF no longer has editable vectors or form fields

Cause: The output is reconstructed from rendered page images plus invisible OCR text; it does not copy the source document structure.

Fix: Keep the source. Use a text-layer extractor when the existing selectable text already meets the task, and inspect the new PDF before replacing anything.

A confident result contains a wrong amount or reading order

Cause: Confidence is an engine estimate; tables, columns, mixed language and poor scans can still be misread.

Fix: Choose the matching language and layout, compare digits with the source and try a higher scale within the pixel budget. OCR is not a table-to-spreadsheet converter.

OCR stops before completion or finds no text

Cause: Language assets may fail to load, a page can exceed the render budget, or scans may contain no recognizable text.

Fix: Retry asset loading, select fewer pages or lower the scale for budget errors. Use clearer scans or another language for empty results. Partial results are not exported.

Production Snippets

A resolution change keeps the same physical page

text

Source displayed size: 800 × 250 pt
1.5× render: 1200 × 375 pixels (108 dpi)
2.5× render: 2000 × 625 pixels (180 dpi)
Both searchable outputs: 800 × 250 pt
Expected example text includes: Invoice 2026-0930 / Total USD 128.50

Frequently Asked Questions

How does OCR differ from text-layer extraction?

It recognizes rendered pixels, including scans with no text layer. Existing selectable text is also rasterized; use extraction when recognition is unnecessary.

What survives in the searchable PDF?

Rendered images and a new invisible OCR text layer at the original physical displayed sizes. Original vectors, forms, annotations, links, bookmarks and signatures are not retained.

Does a higher render scale enlarge the paper?

No. Image and invisible text are scaled together to the source displayed physical size. Higher scale raises pixel resolution and memory usage only.

Which languages and layouts are available?

English, Simplified Chinese or both, with automatic layout, a single text block or sparse text. Confidence is an OCR estimate; check names, numbers and reading order.

What limits apply to selection and output?

75 MiB input, 1000 source pages, at most 40 selected pages, 16 million pixels per page and 80 million total. PDF output is limited to 100 MiB; TXT/JSON to 8 MiB. Invalid ranges reject the whole job.

Are files uploaded or partial results exported?

PDFs stay local; only engine and language assets download. A failed or all-empty job has no completed export. File/option changes or leaving the page invalidate work, workers and old results; downloads are explicit.

Keep browsing