How DOCX Files Store Their Content
A .docx file is a ZIP archive of XML parts. The actual text lives in word/document.xml, stored as a tree of paragraphs and “runs” of characters; images sit in word/media/, and formatting is kept separately in word/styles.xml and referenced by the text. Like .xlsx and .pptx, it’s the Office Open XML format — which is exactly why a Word document can be read and extracted without Word. Here’s how the pieces fit together.
This guide opens up the structure of a DOCX and shows how to pull the words out with the DOCX text extractor.
A DOCX is a ZIP — look inside
Copy any .docx, rename it to .zip, and open it. You’ll find a small tree of folders and XML files rather than one binary blob. This packaging is the Open Packaging Convention, and the XML inside is Office Open XML (OOXML), the same family Excel and PowerPoint use.
The main parts
word/document.xml— the document body: all the paragraphs and text. This is where the readable content lives.word/styles.xml— the style definitions (Heading 1, Normal, fonts, spacing), referenced by the body rather than repeated in it.word/media/— the embedded images, each stored as a normal image file (image1.png, etc.).word/_rels/document.xml.rels— relationships linking the body to its images, headers and hyperlinks.docProps/core.xmlandapp.xml— the document metadata (author, dates, editing time).
How the text is actually stored: paragraphs and runs
Inside document.xml, text is organised as a hierarchy. A paragraph is a <w:p> element. Within it, text is split into runs — <w:r> elements — where each run is a stretch of characters sharing the same formatting, and the characters themselves sit in a <w:t> (text) element:
<w:p>
<w:r><w:t>The quick </w:t></w:r>
<w:r><w:rPr><w:b/></w:rPr><w:t>brown</w:t></w:r>
<w:r><w:t> fox</w:t></w:r>
</w:p>
That’s a single sentence — “The quick brown fox” — split into three runs because the middle word is bold. This is why you can’t just read the XML and get clean prose: a text extractor has to walk the paragraphs, concatenate the runs in order, and drop the formatting markup to reconstruct the readable text.
What this means for extraction
Because the format is open XML, extracting a Word document’s text is reliable and needs no Word install. The DOCX text extractor reads document.xml, joins the runs back into paragraphs, and gives you clean text in your browser, with the file never uploaded. If you want the pictures instead of the words, the DOCX image extractor pulls everything out of word/media/ at full quality.
Frequently asked questions
How does a DOCX file store text?
In word/document.xml, as a tree of paragraphs (<w:p>) containing runs (<w:r>) of characters. Each run is a stretch of text sharing one format.
Is a DOCX file really a ZIP?
Yes. A .docx is a ZIP archive of XML parts (Office Open XML). Rename a copy to .zip and you can browse the document body, styles and images inside.
Where are images stored in a Word document?
In the word/media/ folder inside the archive, each as an ordinary image file. A DOCX image extractor pulls them out at full quality.
Why is a sentence split into several runs?
Because formatting changes within it. Each run holds characters with one consistent format, so a bold word in the middle of a sentence starts a new run.
Can I read a DOCX without Microsoft Word?
Yes. Because the format is open XML, tools like the DOCX text extractor read it directly in the browser.
Related reading
Last updated: 16 August 2026.