Skip to content
UniKit

HTML content extractor

Paste HTML and extract every link (text, href, external and nofollow flags), images, the h1–h6 heading tree, meta tags, plain text and table data — copy per category or export JSON.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

HTML source
Extracted content

Paste some HTML first

JSON export

What this tool does

  • Audit a page in one pass: every link, image, heading and meta tag in one place, with external and nofollow flags already marked.
  • Pull tables and body text out of a saved HTML file so you can reuse the data as plain text or JSON.
  • Check the heading outline for skipped levels or multiple h1s — paste the markup and read the indented tree.
  • Get just the readable copy of a page by stripping tags and scripts first, then use it for word counts or translation.

Example

Input

<h1>UniKit guide</h1>
<p>Open <a href="/tools">the tool list</a>, or visit <a href="https://astro.build" rel="nofollow" target="_blank">Astro</a>.</p>
<img src="/cover.png" alt="cover" width="800" height="400">

Output

the tool list	/tools
Astro	https://astro.build [external, nofollow, _blank]

(headings) h1 UniKit guide
(stats) links 2 · images 1 · headings 1 · tables 0 · characters 40 · words 9

The "external" flag is relative to the base URL: with https://unikit.cc/guide set, /tools counts as internal while https://astro.build is external. Leave it empty and every absolute URL is treated as external.

Frequently asked questions

What happens if I leave the base URL empty?

Relative hrefs such as /tools or ./a.html are never flagged as external, while every absolute URL (http://, https:// or //) is. Fill in the page URL to get a real comparison against the host name.

Does the HTML have to be well-formed?

No. The parser is a forgiving hand-written tokenizer: a missing closing tag simply lets the following content continue as children instead of failing. The only hard error is empty input.

Why is the text inside script and style tags missing from the plain text output?

Because those blocks are code rather than page copy, they are dropped entirely. Table cells are joined with a space, block elements add line breaks, and empty lines are removed at the end.

How is the word count computed?

Latin words are counted as words and CJK characters are counted individually, then the two are added together. That is why a mixed Chinese/English page reports a larger count than the English text alone would suggest.

Is the HTML I paste uploaded anywhere?

No. Parsing, extraction and the JSON export all run in your browser; the page makes no network requests and your markup is never logged.

Keywords:html 提取html extractor链接提取extract links标题层级heading treemeta 提取表格提取table extractorhtml 解析

Related tools