HTML code cleaner
Content pasted from Word, exported from a page builder or scraped from a live site arrives full of junk. Formatting it is how you find the junk.
Cleaning is a two-step job
First make the mess visible, then remove it deliberately. A tool that guesses which tags are disposable will eventually guess wrong on the one that mattered.
Where dirty HTML comes from
WYSIWYG editors. Every time someone bolds a word, deletes it and types something else, the
editor can leave an empty <strong> or a <span> with no attributes.
Do that for a year and a 400-word article carries a thousand tags.
Pasting from Word or Google Docs. Office formats carry their own markup, and browsers
translate it into HTML on paste. The result is class="MsoNormal", inline
style="mso-..." declarations and nested <font> tags that no stylesheet will
ever match.
Page builders and email tools. Drag-and-drop builders wrap every element in two or three containers so they have somewhere to attach handles and spacing controls. That scaffolding ships to production.
Scrapers and archives. Saved pages come back with injected analytics, absolute URLs pointing at the original domain and, occasionally, comment blocks from the tool that saved them.
Cleaning in two passes
Format it so you can see it
Paste into the panel above and press Format. The indentation exposes the structure: rows of identical
nested wrappers, empty elements and the stray <br> tags that were invisible in a wall
of text.
Check what the structure is hiding
Switch to Validate. Duplicate IDs from copied blocks, unclosed tags from a truncated paste and images
with no alt all surface with line numbers.
Delete by hand, then reformat
The removals are judgement calls — only you know whether that wrapper is load-bearing. Edit in the left panel, press Format again, and repeat until the structure matches what the content actually needs.
A checklist for pasted content
- Remove
class="Mso*"and anystyleattribute starting withmso-. - Delete empty elements:
<span></span>,<p> </p>,<strong></strong>. - Replace
<b>and<i>with<strong>and<em>where the emphasis is meaningful, or drop them where it is decorative. - Replace runs of
<br><br>with real paragraphs. - Strip inline
styleattributes and let your stylesheet do the work. - Check heading order — pasted content often arrives with an
<h1>that should be an<h2>, or jumps from<h2>to<h4>. - Make sure links that opened in a new tab still carry
rel="noopener".
Questions about this tool
Does this remove Microsoft Word formatting automatically?
It does not delete anything on its own. It formats the markup so the Word artefacts are visible and easy to select, and flags structural problems. The deletions are yours to make, because a rule that strips every <span> will eventually strip one you needed.
Can I clean up a whole exported site?
One file at a time here. For a whole site you want a script — html-tidy, a Node script using jsdom, or a find-and-replace across the directory in your editor. Use this page to work out exactly what needs removing on one representative file first.
Will cleaning break my page?
It can, if you remove a wrapper that a CSS selector depends on. Work on a copy, remove one category of junk at a time, and check the rendered page between passes.
Related tools
HTML Formatter
Indent and structure any HTML file.
OpenHTML Beautifier
Turn minified HTML back into readable code.
OpenHTML Pretty Printer
Pretty-print HTML with the indent you want.
OpenHTML Minifier
Strip whitespace and comments from HTML.
OpenHTML Validator
Find unclosed tags and broken nesting.
OpenHTML to JSX
Convert HTML into React-ready JSX.
OpenLearn the why, not just the how
Longer reading on formatting, indentation and minification.