Skip to content
UniKit

HTML tag stripper

Strip every tag out of an HTML snippet and keep only the text: optionally keep line breaks and link targets, decode entities, drop script/style bodies and count how many tags were removed.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

HTML source

Options

Plain text

What this tool does

  • Copy the body of a web page and strip the markup to get plain text you can paste into a document, an email or a spreadsheet.
  • Clean up exported HTML mail or scraped pages: the whole <script>/<style> blocks are dropped so code never leaks into the copy.
  • With links kept, each anchor text is followed by its URL (text <https://…>), so you can collect references without opening every link.
  • Turn HTML tables or lists into newline-separated text that Excel or a shell pipeline can chew on.

Example

Input

<h1>UniKit</h1>
<p>Strip <b>HTML</b> tags but keep the <a href="https://unikit.cc">link</a> and the text.</p>
<script>console.log("x")</script>

Output

UniKit
Strip HTML tags but keep the link <https://unikit.cc> and the text.

With the defaults: block tags become line breaks, the <script> block is removed entirely, entities such as &amp; are decoded and runs of whitespace collapse to one space. The tool also reports 10 tags and 1 script block removed.

Frequently asked questions

Why does the link text get a <https://…> appended?

That is the "keep links" option: it appends the href after the anchor text so references stay usable. Turn it off and the URL disappears with the tag. Pure fragment links starting with # never get an address appended.

How are table rows separated in the output?

td, th, tr and table are all block tags that convert to line breaks, so every cell and row starts on its own line. If you want one record per line, keep whitespace collapsing on so consecutive blank lines merge.

What if the HTML is messy and tags are not closed?

Nothing breaks. The tokenizer is deliberately lenient: a stray < is treated as plain text and unclosed tags never abort the run. The only hard limit is size — over 2 million characters and the tool asks you to split the input.

How is this different from the HTML content extractor?

The extractor is built for analysis: link and image tables, a heading tree, meta tags, two-dimensional table data and a JSON export. This tool just emits one block of plain text, which makes it the faster choice when you only need the copy.

Is the text I paste uploaded?

No. Everything is processed in browser memory, the page makes no network requests, and nothing you paste is stored.

Keywords:strip htmlhtml to text去掉 html 标签html 转文本html tag stripperplain text纯文本解码实体decode entitiesdom to text

Related tools