Why stripping HTML properly is harder than it looks
The obvious approach — delete everything between < and > with a
find-and-replace — fails on real pages in three predictable ways. It leaves the contents of
<script> and <style> tags in the output, so you get lines of
JavaScript and CSS mixed into the text. It leaves entities like & and
as literal character sequences. And it runs every block together, because the tags
that created the line breaks are exactly what got deleted.
This tool parses the HTML with the browser’s own HTML parser instead — the same one a
browser uses to render a page. That means <script> and <style> are
identified and dropped, entities are decoded to real characters, and the document structure is known, so a
line break can be kept wherever a paragraph, heading or list item ended.
How to strip HTML to plain text
- Paste the HTML. Drop the markup into the box above — a page’s View Source, an HTML email, a fragment copied from a template. It does not need to be a complete document.
- Choose how lists and links are handled. By default list items are marked with a dash so structure survives. Turn on Keep link addresses to append each link’s URL in angle brackets after its text.
- Read the result. The text comes back with scripts, styles and hidden markup gone, entities decoded, and a line break where each paragraph, heading or list item was.
- Copy or save. Copy text puts the result on your clipboard; Save as .txt writes it to a file.
What comes out
- The readable text, with markup removed and entities decoded to real characters.
- Preserved structure — a blank line between paragraphs, a line break after each heading, and each list item on its own line.
- A word count of the result.
Two options adjust the output. Mark list items prefixes each with a dash so a list still reads as a list. Keep link addresses appends each link’s URL in angle brackets, which turns the text into something you can still follow the links from — useful when the links are the point.
What is handled, and what is not
Handled: complete HTML documents and fragments alike. <script>,
<style>, <noscript>, <template>,
<svg> and the document head are dropped. Every HTML entity is decoded. Block elements
— paragraphs, headings, list items, table rows, <br> — produce line breaks;
runs of whitespace are collapsed the way a browser would, so the output is not full of the indentation
from the source.
Not done: this returns text, not Markdown — it does not reproduce bold, italics or heading levels as symbols. Tables are flattened to their cell text rather than redrawn; for tabular data specifically, the HTML table extractor turns a table into CSV with its columns intact. Parsing is inert: the HTML is never rendered and no script in it runs, so pasting untrusted markup is safe.
Why the HTML is never uploaded
The markup is parsed and stripped in your browser. No server takes part, so nothing is transmitted.
The parser used, DOMParser, builds the document without rendering it — no script
runs and no image, stylesheet or tracker referenced in the HTML is ever fetched. That makes it safe to
paste markup from an email or a page you do not trust: reading its text here contacts none of the servers
that opening it in a browser would.
What this tool does not do
Three boundaries:
- It does not fetch a page from a URL. Paste the HTML in. To get a page’s source, use your browser’s View Source or Save As.
- It does not convert to Markdown. The output is plain text; formatting like bold and headings becomes plain words, not symbols.
- It does not keep table layout. A table’s cells come out as text. For a table as CSV, use the HTML table extractor.
Why people strip HTML to text
- Getting the words out of an email whose source is full of markup and inline styles.
- Cleaning copied content — pasting a chunk of a page and keeping only the readable text.
- Preparing text for analysis — feeding clean prose into a word counter, a summariser or a search rather than raw markup.
- Reading a suspicious HTML email safely — seeing its text without rendering it or loading anything it references.
- Extracting content from a template without the surrounding structure.
HTML to text compared with the other HTML tools
This tool gives you the readable prose. When you want a specific part of the HTML instead, a narrower tool keeps its structure: the HTML table extractor turns a table into CSV, and the URL extractor pulls every link out of the markup and groups them by domain.
For text that is inside a document rather than in HTML, use the document text extractors or the PDF text extractor. To match a specific pattern in the text once you have it, use the regex extractor.
HTML text format and extraction edge cases
HTML text content is the text found in DOM text nodes — the characters between
element tags. Inline elements (<strong>, <em>,
<a>, <span>) contribute their text in flow; block
elements (<p>, <div>,
<h1>–<h6>, <li>) each generate a
line break in the output.
Four edge cases: (a) <script> and <style> element
content is excluded — their text nodes are code, not readable prose; (b) HTML entities
(& < ) are decoded to their
Unicode characters before output; (c) non-breaking spaces ( from
) are replaced with ordinary spaces to prevent invisible
word-boundary issues in downstream text processing; (d) per the CSS whitespace-collapsing
rules, consecutive spaces and tabs in flow content collapse to a single space, matching what
is visible in the browser.
Frequently asked questions
How do I convert HTML to plain text?
Paste the HTML into the box above and click Strip to text. Tags are removed, entities are decoded, and paragraphs and lists keep their line breaks.
Why not just delete everything between the angle brackets?
Because that leaves the contents of script and style tags in the output, leaves entities like & as literal text, and runs every paragraph together. Parsing the HTML properly avoids all three.
Does it run the JavaScript in the HTML?
No. The markup is parsed but never rendered, so no script runs and nothing the HTML references is fetched. That is what makes it safe for untrusted markup.
Can it keep the links?
Yes. Turn on Keep link addresses and each link’s URL is added in angle brackets after its text, so you can still follow it.
Does it turn HTML into Markdown?
No. The output is plain text. Formatting such as bold and headings becomes plain words rather than Markdown symbols.
Can I paste a whole web page?
Yes. Paste the full HTML source and the readable text comes back. It handles complete documents and small fragments alike.
Is my HTML uploaded?
No. It is parsed and stripped inside your browser, and nothing is transmitted.