HTML to Text
Strip markup and retain selected text breaks
Paste HTML to remove tags and complete script/style blocks, with selected line breaks. The output omits href destinations and does not reconstruct tables.
The full guide also includes pitfalls, worked examples, snippets, FAQs, and related tools for checking results or troubleshooting.
About this tool
Convert pasted markup to plain text by removing tags and complete script/style blocks. Selected block endings and br tags become line breaks, and li items receive bullet markers. Six common entity spellings are decoded. This is lightweight text cleanup: it does not preserve link destinations, infer visible page content from CSS, reconstruct table columns, or implement the complete HTML entity set.
Production Snippets
Line breaks and supported entities differ from full HTML decoding
text
INPUT
Hello<br>World<br>& A<script>skip()</script>
OUTPUT
Hello
World
& A
& is decoded; the hexadecimal entity A stays literal. This converter does not implement the complete HTML entity set.Frequently Asked Questions
Which line breaks are retained?
br and closing p, div, li, and h1–h6 tags create newlines; opening li adds a bullet. Other block tags and table cells are not structurally reconstructed, and some spaces around line breaks may remain.
Which HTML entities are decoded?
Only , &, <, >, ', and " are explicitly replaced. Other named or numeric entities remain literal. The sequence of replacements is not a complete HTML entity parser.
What happens to links, hidden content, and scripts?
Link labels remain as text while href values are discarded. Complete script/style blocks are removed. CSS visibility and hidden attributes are not evaluated, so hidden text can still appear.
Is the output sanitized HTML or a full page transcript?
No. The result is plain text shown as text, not an HTML security sanitizer or browser layout reconstruction. Do not insert it as trusted HTML; use a structured parser when exact document fidelity is required.
Keep browsing