TXT

Extract PDF Text

Export an existing PDF text layer as TXT or JSON

Render & Extract
πŸ”’ 100% client-side β€” your data never leaves this page
Maintained by Evanβ€’Updated: September 30, 2026
⏳Loading tool…

About this tool

Extract the existing text layer from selected pages without running OCR. Source item order and explicit line-ending signals produce TXT; JSON also retains original item strings, coordinates, direction and page positions. This is not visual layout reconstruction: columns, rotations, ligatures and PDF.js spacing normalization can affect reading order. A strict range such as 1-3,5 rejects invalid tokens and sorts duplicate selections into source order. All-empty pages produce a no-text message rather than a file of page headings. CMap or font-resource failures stop the operation. The preview is limited to 16,000 characters, while explicit TXT/JSON downloads contain the full bounded result. Limits: 20 MiB, 200 pages, 4 MiB of extracted text and 100,000 items.

Suggested Workflow

Scenario Recipes

01

Export a selectable invoice text layer

Goal: Export an existing PDF text layer as TXT or JSON

  1. Load the example and enter page range 1.
  2. Extract text and inspect Invoice 128.50 in the preview.
  3. Download TXT for plain text or JSON for the page and item coordinates.

Result: The output contains actual text, not an OCR claim or page headings alone.

Failure Clinic (Common Pitfalls)

Will this recognize a scanned page?

Cause: No. A page without a usable text layer needs OCR first. An empty extraction creates no completed text file.

Fix: Use the PDF OCR tool for a scan, then extract the resulting searchable text layer.

Does TXT preserve the visual layout?

Cause: No. It preserves PDF.js source item order and explicit line endings. Use JSON coordinates or a dedicated layout workflow for columns and rotated text.

Fix: Use item transforms in JSON to build a layout-specific reader; check rotated and multi-column pages manually.

Production Snippets

Export a selectable invoice text layer

text

Range: 1
--- Page 1 ---
Invoice 128.50
Selectable sample text.

Frequently Asked Questions

Will this recognize a scanned page?

No. A page without a usable text layer needs OCR first. An empty extraction creates no completed text file.

Does TXT preserve the visual layout?

No. It preserves PDF.js source item order and explicit line endings. Use JSON coordinates or a dedicated layout workflow for columns and rotated text.

Why did a partly valid page range fail?

Every token must be a complete in-range page number or ascending range. Invalid tokens are not silently discarded.

Keep browsing