H2T

HTML to Text

Strip markup and retain selected text breaks

Document & Media
🔒 100% client-side — your data never leaves this page
Maintained by Evan•Updated: September 29, 2026
Options
HTML Input

Paste HTML to remove tags and complete script/style blocks, with selected line breaks. The output omits href destinations and does not reconstruct tables.

Plain Text
Plain text will appear here
🔒 Processed locally in your browser
Page reading mode

The full guide also includes pitfalls, worked examples, snippets, FAQs, and related tools for checking results or troubleshooting.

About this tool

Convert pasted markup to plain text by removing tags and complete script/style blocks. Selected block endings and br tags become line breaks, and li items receive bullet markers. Six common entity spellings are decoded. This is lightweight text cleanup: it does not preserve link destinations, infer visible page content from CSS, reconstruct table columns, or implement the complete HTML entity set.

Production Snippets

Line breaks and supported entities differ from full HTML decoding

text

INPUT
Hello<br>World<br>&amp; &#x41;<script>skip()</script>

OUTPUT
Hello
World
& &#x41;

&amp; is decoded; the hexadecimal entity &#x41; stays literal. This converter does not implement the complete HTML entity set.

Frequently Asked Questions

Which line breaks are retained?

br and closing p, div, li, and h1–h6 tags create newlines; opening li adds a bullet. Other block tags and table cells are not structurally reconstructed, and some spaces around line breaks may remain.

Which HTML entities are decoded?

Only &nbsp;, &amp;, &lt;, &gt;, &#39;, and &quot; are explicitly replaced. Other named or numeric entities remain literal. The sequence of replacements is not a complete HTML entity parser.

What happens to links, hidden content, and scripts?

Link labels remain as text while href values are discarded. Complete script/style blocks are removed. CSS visibility and hidden attributes are not evaluated, so hidden text can still appear.

Is the output sanitized HTML or a full page transcript?

No. The result is plain text shown as text, not an HTML security sanitizer or browser layout reconstruction. Do not insert it as trusted HTML; use a structured parser when exact document fidelity is required.

Keep browsing