Extract Text From a Word Document

A .docx stores its text as XML inside a ZIP, with headings and lists recorded as styles rather than as characters โ€” which is why pasting out of Word loses the structure. Drop a document below to get its text back as clean plain text, or as Markdown that keeps the headings, lists, tables, bold, italics and links. The file is read in your browser and never uploaded.

Drop a .docx here
or click to choose a file · nothing is uploaded

What is actually inside a .docx file

A .docx is a ZIP archive of XML parts. The text lives in word/document.xml as a sequence of paragraphs, each containing runs of characters. A heading is not big bold text in the file โ€” it is an ordinary paragraph carrying a style reference of Heading1. A bullet is a paragraph with numbering properties attached.

That separation is why copying out of Word so often disappoints. The visual structure lives in the styles, and a plain copy takes only the characters. Extraction that reads the styles as well can rebuild the structure, which is what the Markdown output here does.

For a full technical breakdown of every part inside a .docx archive, read how DOCX files store content.

How to extract text from a DOCX file

  1. Open the document. Drop the .docx onto the box above, or click to browse. Only the document's XML is read, so a file full of images opens as quickly as a plain one.
  2. Choose plain text or Markdown. Plain text gives you the words with no markup. Markdown keeps headings as # levels, lists as bullets, tables as pipe tables, and bold, italics and links as Markdown syntax.
  3. Decide how tables should come out. Keep tables as tables to preserve their grid. Switch it off to flatten every table into tab-separated rows you can paste into a spreadsheet.
  4. Copy or download. Copy the result, or download it as a .txt or .md file. Markdown downloads carry the .md extension so editors highlight it correctly.

What comes out

Which files work

.docx files from Word 2007 onward work, along with .docx exported by Google Docs, LibreOffice, Pages and Office Online. Documents of several hundred pages process in well under a second, because only the XML is parsed and images are skipped entirely.

The legacy .doc format is a different, binary format and is detected and reported rather than mangled โ€” open it in Word or LibreOffice and save as .docx first. Password-protected documents are encrypted at the container level and must be unlocked before extraction.

Why the document is never uploaded

The archive is read from your disk by the browser and decompressed with the browser's own DecompressionStream. There is no server step, so the document is never transmitted and no copy exists after the tab closes.

Word documents are among the most sensitive files people convert online: contracts, offer letters, medical notes, legal drafts. Every "docx to text" service that asks you to upload has a copy of that document on its infrastructure, often for hours. Reading it locally removes the question entirely.

What is not preserved

Plain text and Markdown cannot carry everything a Word document holds, so five things are dropped by design:

Text boxes and shapes are also skipped, because their content sits outside the main document flow.

Who extracts text from Word files

DOCX extraction compared with PDF extraction

A DOCX is far easier to extract from than a PDF, and the reason is structural. A DOCX records paragraphs, headings and tables explicitly, so extraction reads intent directly. A PDF records glyphs at coordinates, so extraction has to infer reading order from geometry โ€” which is why the PDF text extractor rebuilds lines from baselines and can interleave columns.

If your document is a PDF, use that tool instead. If it is a scan with no text layer, start with image to text OCR. To see the raw parts inside a DOCX โ€” including the images in word/media/ โ€” open it with the ZIP extractor, since a DOCX is a ZIP container.

DOCX XML structure and extraction edge cases

For a plain-language explanation of the format, see the guide on how DOCX files store content. A DOCX is an OOXML package per ISO/IEC 29500. The document body lives in word/document.xml. Paragraphs are <w:p> elements; text runs are <w:r><w:t>. Part relationships are declared in _rels/ companion directories. Heading levels come from the style reference (<w:pStyle w:val="Heading1"/>), not from visual font size.

Three edge cases that affect what comes out: (a) text boxes โ€” <w:drawing><wp:inline><w:txbxContent> โ€” sit outside the main document body flow and are skipped, so captions inside floating text boxes do not appear in the output; (b) revision tracking stores both deleted (<w:del>) and inserted (<w:ins>) runs โ€” only the accepted body is returned, meaning deleted runs are absent; (c) field results (<w:fldSimple> or <w:instrText>) store a cached calculated value such as a date or cross-reference โ€” the extractor returns the cached result, which was current when the document was last saved but may be stale if fields were not updated before closing.

Frequently asked questions

How do I extract text from a DOCX without Word?

Drop the file onto this page. It is read in your browser, so no Word licence and no software install is needed.

Can I convert a Word document to Markdown?

Yes. Choose Markdown and the output keeps headings as # levels, lists as bullets, tables as pipe tables, and bold, italics and links as Markdown syntax.

Is my document uploaded to a server?

No. The .docx is read from your disk and unzipped by your browser. It is never transmitted.

Does it keep my tables?

Yes. Tables come out as Markdown pipe tables, or as tab-separated rows if you switch off Keep tables as tables โ€” which is the better choice for pasting into a spreadsheet.

Why does it refuse my .doc file?

The old .doc format is binary and completely different from .docx. Open it in Word or LibreOffice and save it as .docx, then extract.

Are images extracted too?

Not by this tool โ€” images are counted and reported. Because a .docx is a ZIP archive, you can open it with the ZIP extractor and save anything under word/media/ directly.

What about headers, footers and footnotes?

Those are stored in separate XML parts and are not included in the body text. Only the main document flow is returned.

Does it work on Mac, Linux or a phone?

Yes. Everything runs in the browser, so no operating system or Office install is required.

• Specialist file parsing & security engineer • Verified: in our experience, our hands-on testing measured and verified private in-browser execution with zero file uploads • Last reviewed July 2026.