Skip to content
UniKit

HTML to Markdown

Convert HTML to Markdown with a hand-written parser: headings h1–h6, paragraphs, bold / italic / strikethrough, inline code and fenced code blocks with a language, nested lists, links, images, blockquotes, rules and GFM tables, with HTML entities decoded.

Runs in your browserEvery computation happens in your browser — your data never leaves this device.

Markdown

Type some HTML above and the result shows up here.

What this tool does

  • Turn HTML copied from a web page, CMS or rich-text editor into Markdown you can paste into a README, wiki or blog draft.
  • Migrate a documentation site: convert a batch of HTML pages into Markdown sources for a static site generator.
  • Clean up rich text from an email or ticket — script and style noise is dropped and only readable content remains.
  • Keep structure intact instead of flattening it: GFM tables, nested lists and fenced code blocks with a language all survive the conversion.

Example

Input

<h1>Title</h1><p>Body <strong>bold</strong> and <a href="https://a.com">link</a></p>

Output

# Title

Body **bold** and [link](https://a.com)

script / style / head content is dropped, and code inside pre is carried over verbatim into a fenced block, keeping its language-xxx marker.

Frequently asked questions

Is the output CommonMark or GFM?

CommonMark is the baseline; tables use the GitHub Flavored Markdown syntax (`| a | b |` plus an alignment row) and strikethrough uses `~~text~~`. Both render as-is on GitHub, GitLab and most static site generators.

Are nested lists indented correctly?

Yes. Each nesting level is indented by the width of its parent marker, and ordered lists honour `<ol start="3">`, producing `3.` instead of `1.`.

Where does the code block language come from?

From the class name in `<pre><code class="language-xxx">`, which becomes a ```xxx fence; without a class you get a plain fence. HTML entities inside the code are decoded, but the code itself is neither escaped nor rewritten.

Are * _ [ ] escaped in the body text?

Yes. Inside text nodes, `\`, backticks, `*`, `_`, `[` and `]` get a backslash so they are not read as Markdown syntax, and `|` is escaped inside table cells so the table does not fall apart.

Do unclosed tags raise an error?

No. HTML allows optional closing tags such as `</p>` and `</li>`, so they are closed automatically. Only genuinely truncated input — a tag without `>` or a comment without `-->` — is reported as an error.

Keywords:html to markdownhtml2mdmarkdown converterhtml 转 markdownhtml 转 md富文本转换gfm 表格代码块嵌套列表实体解码文档转换

Related tools