What is actually inside a .docx file
A .docx is a ZIP archive of XML parts. The text lives in
word/document.xml as a sequence of paragraphs, each containing runs of characters. A heading
is not big bold text in the file โ it is an ordinary paragraph carrying a style reference of
Heading1. A bullet is a paragraph with numbering properties attached.
That separation is why copying out of Word so often disappoints. The visual structure lives in the styles, and a plain copy takes only the characters. Extraction that reads the styles as well can rebuild the structure, which is what the Markdown output here does.
For a full technical breakdown of every part inside a .docx archive, read how DOCX files store content.
How to extract text from a DOCX file
- Open the document. Drop the .docx onto the box above, or click to browse. Only the document's XML is read, so a file full of images opens as quickly as a plain one.
- Choose plain text or Markdown. Plain text gives you the words with no markup. Markdown keeps headings as # levels, lists as bullets, tables as pipe tables, and bold, italics and links as Markdown syntax.
- Decide how tables should come out. Keep tables as tables to preserve their grid. Switch it off to flatten every table into tab-separated rows you can paste into a spreadsheet.
- Copy or download. Copy the result, or download it as a .txt or .md file. Markdown downloads carry the .md extension so editors highlight it correctly.
What comes out
- All body text in document order, including text inside tables.
- Headings at their real levels, one to six, as Markdown
#prefixes. - Lists with their nesting preserved, indented by level.
- Tables either as Markdown pipe tables or as tab-separated rows for pasting into Excel or Google Sheets.
- Bold, italic and hyperlinks, with link targets resolved from the document's relationships so you get the real URL rather than the display text.
- Document properties โ title, author, last modified date and the application that produced the file.
- Counts โ paragraphs, tables, words and how many images the file contains.
Which files work
.docx files from Word 2007 onward work, along with .docx exported by Google
Docs, LibreOffice, Pages and Office Online. Documents of several hundred pages process in well under a
second, because only the XML is parsed and images are skipped entirely.
The legacy .doc format is a different, binary format and is detected and reported rather
than mangled โ open it in Word or LibreOffice and save as .docx first. Password-protected
documents are encrypted at the container level and must be unlocked before extraction.
Why the document is never uploaded
The archive is read from your disk by the browser and decompressed with the browser's own DecompressionStream. There is no server step, so the document is never transmitted and no copy exists after the tab closes.
Word documents are among the most sensitive files people convert online: contracts, offer letters, medical notes, legal drafts. Every "docx to text" service that asks you to upload has a copy of that document on its infrastructure, often for hours. Reading it locally removes the question entirely.
What is not preserved
Plain text and Markdown cannot carry everything a Word document holds, so five things are dropped by design:
- Images. They are counted and reported, but not extracted here.
- Headers, footers and footnotes. These live in separate XML parts and are not included in the body text.
- Comments and tracked changes. Only the accepted document text is returned.
- Fonts, colours, sizes and spacing. Markdown records structure, not appearance.
- Merged table cells. A merged cell becomes one cell in its row, so the grid can be narrower than it looks on screen.
Text boxes and shapes are also skipped, because their content sits outside the main document flow.
Who extracts text from Word files
- Writers moving to Markdown โ converting a Word draft into a format a static site generator, wiki or documentation tool accepts.
- Developers โ getting document content into a repository, changelog or CMS without Word's markup coming along.
- Anyone feeding text to an AI tool โ producing clean text to paste into a model that cannot read .docx directly.
- Migration work โ bulk-converting documents into a system that only accepts plain text.
- People without Word โ reading a document on a machine or phone with no Office licence installed.
DOCX extraction compared with PDF extraction
A DOCX is far easier to extract from than a PDF, and the reason is structural. A DOCX records paragraphs, headings and tables explicitly, so extraction reads intent directly. A PDF records glyphs at coordinates, so extraction has to infer reading order from geometry โ which is why the PDF text extractor rebuilds lines from baselines and can interleave columns.
If your document is a PDF, use that tool instead. If it is a scan with no text layer, start with
image to text OCR. To see the raw parts inside a DOCX โ including the
images in word/media/ โ open it with the ZIP extractor, since
a DOCX is a ZIP container.
DOCX XML structure and extraction edge cases
For a plain-language explanation of the format, see the guide on how DOCX files store content. A DOCX is an OOXML package per ISO/IEC 29500. The document body lives in
word/document.xml. Paragraphs are <w:p> elements; text runs
are <w:r><w:t>. Part relationships are declared in
_rels/ companion directories. Heading levels come from the style reference
(<w:pStyle w:val="Heading1"/>), not from visual font size.
Three edge cases that affect what comes out: (a) text boxes โ
<w:drawing><wp:inline><w:txbxContent> โ sit outside the
main document body flow and are skipped, so captions inside floating text boxes do not appear
in the output; (b) revision tracking stores both deleted (<w:del>) and
inserted (<w:ins>) runs โ only the accepted body is returned, meaning
deleted runs are absent; (c) field results (<w:fldSimple> or
<w:instrText>) store a cached calculated value such as a date or
cross-reference โ the extractor returns the cached result, which was current when the
document was last saved but may be stale if fields were not updated before closing.
Frequently asked questions
How do I extract text from a DOCX without Word?
Drop the file onto this page. It is read in your browser, so no Word licence and no software install is needed.
Can I convert a Word document to Markdown?
Yes. Choose Markdown and the output keeps headings as # levels, lists as bullets, tables as pipe tables, and bold, italics and links as Markdown syntax.
Is my document uploaded to a server?
No. The .docx is read from your disk and unzipped by your browser. It is never transmitted.
Does it keep my tables?
Yes. Tables come out as Markdown pipe tables, or as tab-separated rows if you switch off Keep tables as tables โ which is the better choice for pasting into a spreadsheet.
Why does it refuse my .doc file?
The old .doc format is binary and completely different from .docx. Open it in Word or LibreOffice and save it as .docx, then extract.
Are images extracted too?
Not by this tool โ images are counted and reported. Because a .docx is a ZIP archive, you can open it with the ZIP extractor and save anything under word/media/ directly.
What about headers, footers and footnotes?
Those are stored in separate XML parts and are not included in the body text. Only the main document flow is returned.
Does it work on Mac, Linux or a phone?
Yes. Everything runs in the browser, so no operating system or Office install is required.