Compare PDF Text
Compare extracted text and page presence at matching positions
About this tool
Compare extracted source-order text at the same page positions in PDFs A and B. A missing page differs from an existing page with empty text, and the JSON retains both existence flags and null versus empty strings. The optional whitespace mode removes JavaScript whitespace from the extracted strings before comparing; PDF.js may already normalize some source spacing. Two scans with no text receive a no-text-to-compare result, while page-presence differences are still reported. Text matches do not establish visual, layout, image or signature equality. Replacing either input or changing an option clears the prior result. Full JSON captures both input names and the chosen mode. Limits: 20 MiB and 200 pages per PDF; 4 MiB text and 100,000 items per input.
Scenario Recipes
Distinguish an extra blank page
Goal: Compare extracted text and page presence at matching positions
- Choose a one-page blank PDF as A and a two-page blank PDF as B.
- Compare text and inspect the page-presence difference count.
- Download JSON and inspect existsA=false on page 2.
Result: There is no text to compare, but one page position differs.
Failure Clinic (Common Pitfalls)
Do two image-only PDFs compare as identical?
Cause: No. Without extractable text the result says there is no text to compare. Extra or missing page positions are still reported.
Fix: Run OCR first if the question concerns scanned text; use a visual comparison tool for image or layout differences.
What does ignore whitespace remove?
Cause: It removes JavaScript whitespace characters from the already extracted text. It does not reconstruct source spacing lost during PDF text extraction.
Fix: Turn the option off when extracted spaces matter and inspect the textA/textB strings in the JSON.
Production Snippets
Distinguish an extra blank page
text
{"page":2,"existsA":false,"existsB":true,"textA":null,"textB":"","equal":false}Frequently Asked Questions
Do two image-only PDFs compare as identical?
No. Without extractable text the result says there is no text to compare. Extra or missing page positions are still reported.
What does ignore whitespace remove?
It removes JavaScript whitespace characters from the already extracted text. It does not reconstruct source spacing lost during PDF text extraction.
Are missing and blank pages different?
Yes. A missing page has exists=false and text=null; a present blank page has exists=true and an empty text string.
Keep browsing