How to convert a Word document to text or HTML
Read a .docx without Word, or strip a document down to text you can actually paste somewhere.
A .docx is a zip containing XML. The words are in there, wrapped in a great deal of formatting information, and this pulls them out.
Step-by-step
- Add your .docx.
- Choose plain text or HTML.
- Copy or download.
Which output to choose
Plain text when you want the words and nothing else — for pasting into a plain-text field, counting, or feeding into another tool.
HTML when structure matters. Headings stay headings, lists stay lists, bold and italic survive. This is the one to use when moving a document into a website or a content system.
Why this beats copy and paste
Copying from Word and pasting into a web editor brings a mass of inline styling with it — fonts, sizes, colours and spacing that fight with the site's own design and are tedious to strip. Converting to clean HTML gives you the structure without the mess.
What does not survive
- Images are not carried into plain text.
- Tracked changes and comments are not part of the readable text.
- Complex layout — text boxes, columns and floated elements — flattens into reading order.
A .docx is a zip full of XML
Modern Word files are not a binary blob. Rename one to .zip, open it, and you will find a folder of XML documents, the images as ordinary files, and a manifest tying them together. The text lives in word/document.xml as marked-up runs.
That is why extraction is reliable and why it can happen entirely in your browser: no rendering engine or reimplementation of Word is required, only reading a zip and walking some XML. It is the same arrangement EPUB uses, and it is the reason both formats can be processed by tools far simpler than the applications that created them.
Deleted text may still be in the file
This is worth knowing before you extract from a document that came from somewhere else. If tracked changes were used and never accepted, the deleted text is still present in the XML, marked as a deletion rather than removed. Comments live in their own part of the archive, complete with author names and timestamps. Neither is visible in the document as most people read it.
So extraction can surface things the sender believed were gone — an earlier price, a name that was taken out, a reviewer's aside. That is occasionally exactly what you want and occasionally an awkward discovery. In the other direction, it is a reason to accept all changes and delete all comments before sending a document to anyone outside, rather than relying on the display to hide them: the display is a view, and the file still has the rest.
Frequently asked questions
Do I need Word installed?
No. The file is unpacked and read in your browser.
Is .doc supported as well as .docx?
The modern .docx format is what this reads. The older binary .doc is a different format entirely; open it in a word processor and save it as .docx first.
Which should I pick for pasting into a website?
HTML. It keeps headings, lists and emphasis without the inline styling that copy-and-paste from Word drags along.
Open the text tools →