{"id":56,"date":"2026-08-19T13:02:00","date_gmt":"2026-08-19T13:02:00","guid":{"rendered":"https:\/\/easyextract.online\/blog\/how-docx-files-store-content\/"},"modified":"2026-08-26T08:14:14","modified_gmt":"2026-08-26T08:14:14","slug":"how-docx-files-store-content","status":"publish","type":"post","link":"https:\/\/easyextract.online\/blog\/how-docx-files-store-content\/","title":{"rendered":"How DOCX Files Store Their Content"},"content":{"rendered":"<p><strong>A <code>.docx<\/code> file is a ZIP archive of XML parts. The actual text lives in <code>word\/document.xml<\/code>, stored as a tree of paragraphs and &#8220;runs&#8221; of characters; images sit in <code>word\/media\/<\/code>, and formatting is kept separately in <code>word\/styles.xml<\/code> and referenced by the text.<\/strong> Like <code>.xlsx<\/code> and <code>.pptx<\/code>, it&#8217;s the Office Open XML format \u2014 which is exactly why a Word document can be read and extracted without Word. Here&#8217;s how the pieces fit together.<\/p>\n<p>This guide opens up the structure of a DOCX and shows how to pull the words out with the <a href=\"https:\/\/easyextract.online\/docx-text-extractor\/\">DOCX text extractor<\/a>.<\/p>\n<h2>A DOCX is a ZIP \u2014 look inside<\/h2>\n<p>Copy any <code>.docx<\/code>, rename it to <code>.zip<\/code>, and open it. You&#8217;ll find a small tree of folders and XML files rather than one binary blob. This packaging is the <strong>Open Packaging Convention<\/strong>, and the XML inside is <strong>Office Open XML (OOXML)<\/strong>, the same family Excel and PowerPoint use.<\/p>\n<h2>The main parts<\/h2>\n<ul>\n<li><strong><code>word\/document.xml<\/code><\/strong> \u2014 the document body: all the paragraphs and text. This is where the readable content lives.<\/li>\n<li><strong><code>word\/styles.xml<\/code><\/strong> \u2014 the style definitions (Heading 1, Normal, fonts, spacing), referenced by the body rather than repeated in it.<\/li>\n<li><strong><code>word\/media\/<\/code><\/strong> \u2014 the embedded images, each stored as a normal image file (<code>image1.png<\/code>, etc.).<\/li>\n<li><strong><code>word\/_rels\/document.xml.rels<\/code><\/strong> \u2014 relationships linking the body to its images, headers and hyperlinks.<\/li>\n<li><strong><code>docProps\/core.xml<\/code> and <code>app.xml<\/code><\/strong> \u2014 the document metadata (author, dates, editing time).<\/li>\n<\/ul>\n<h2>How the text is actually stored: paragraphs and runs<\/h2>\n<p>Inside <code>document.xml<\/code>, text is organised as a hierarchy. A paragraph is a <code>&lt;w:p&gt;<\/code> element. Within it, text is split into <strong>runs<\/strong> \u2014 <code>&lt;w:r&gt;<\/code> elements \u2014 where each run is a stretch of characters sharing the same formatting, and the characters themselves sit in a <code>&lt;w:t&gt;<\/code> (text) element:<\/p>\n<pre><code>&lt;w:p&gt;\n  &lt;w:r&gt;&lt;w:t&gt;The quick &lt;\/w:t&gt;&lt;\/w:r&gt;\n  &lt;w:r&gt;&lt;w:rPr&gt;&lt;w:b\/&gt;&lt;\/w:rPr&gt;&lt;w:t&gt;brown&lt;\/w:t&gt;&lt;\/w:r&gt;\n  &lt;w:r&gt;&lt;w:t&gt; fox&lt;\/w:t&gt;&lt;\/w:r&gt;\n&lt;\/w:p&gt;<\/code><\/pre>\n<p>That&#8217;s a single sentence \u2014 &#8220;The quick <strong>brown<\/strong> fox&#8221; \u2014 split into three runs because the middle word is bold. This is why you can&#8217;t just read the XML and get clean prose: a text extractor has to walk the paragraphs, concatenate the runs in order, and drop the formatting markup to reconstruct the readable text.<\/p>\n<h2>What this means for extraction<\/h2>\n<p>Because the format is open XML, extracting a Word document&#8217;s text is reliable and needs no Word install. The <a href=\"https:\/\/easyextract.online\/docx-text-extractor\/\">DOCX text extractor<\/a> reads <code>document.xml<\/code>, joins the runs back into paragraphs, and gives you clean text in your browser, with the file never uploaded. If you want the pictures instead of the words, the <a href=\"https:\/\/easyextract.online\/docx-image-extractor\/\">DOCX image extractor<\/a> pulls everything out of <code>word\/media\/<\/code> at full quality.<\/p>\n<h2>Frequently asked questions<\/h2>\n<p><strong>How does a DOCX file store text?<\/strong><br \/>\nIn <code>word\/document.xml<\/code>, as a tree of paragraphs (<code>&lt;w:p&gt;<\/code>) containing runs (<code>&lt;w:r&gt;<\/code>) of characters. Each run is a stretch of text sharing one format.<\/p>\n<p><strong>Is a DOCX file really a ZIP?<\/strong><br \/>\nYes. A <code>.docx<\/code> is a ZIP archive of XML parts (Office Open XML). Rename a copy to <code>.zip<\/code> and you can browse the document body, styles and images inside.<\/p>\n<p><strong>Where are images stored in a Word document?<\/strong><br \/>\nIn the <code>word\/media\/<\/code> folder inside the archive, each as an ordinary image file. A <a href=\"https:\/\/easyextract.online\/docx-image-extractor\/\">DOCX image extractor<\/a> pulls them out at full quality.<\/p>\n<p><strong>Why is a sentence split into several runs?<\/strong><br \/>\nBecause formatting changes within it. Each run holds characters with one consistent format, so a bold word in the middle of a sentence starts a new run.<\/p>\n<p><strong>Can I read a DOCX without Microsoft Word?<\/strong><br \/>\nYes. Because the format is open XML, tools like the <a href=\"https:\/\/easyextract.online\/docx-text-extractor\/\">DOCX text extractor<\/a> read it directly in the browser.<\/p>\n<h2>Related reading<\/h2>\n<ul>\n<li><a href=\"https:\/\/easyextract.online\/blog\/what-is-inside-an-xlsx-file\/\">What&#8217;s inside an XLSX file?<\/a><\/li>\n<li><a href=\"https:\/\/easyextract.online\/blog\/hidden-metadata-in-office-files\/\">The hidden metadata in Office files<\/a><\/li>\n<\/ul>\n<p><em>Last updated: 16 August 2026.<\/em><\/p>\n<p><script type=\"application\/ld+json\">\n{\"@context\":\"https:\/\/schema.org\",\"@type\":\"FAQPage\",\"mainEntity\":[\n{\"@type\":\"Question\",\"name\":\"How does a DOCX file store text?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"In word\/document.xml, as a tree of paragraphs containing runs of characters. Each run is a stretch of text sharing one format.\"}},\n{\"@type\":\"Question\",\"name\":\"Is a DOCX file really a ZIP?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Yes. A .docx is a ZIP archive of XML parts (Office Open XML). Rename a copy to .zip and you can browse the document body, styles and images inside.\"}},\n{\"@type\":\"Question\",\"name\":\"Where are images stored in a Word document?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"In the word\/media\/ folder inside the archive, each as an ordinary image file. A DOCX image extractor pulls them out at full quality.\"}},\n{\"@type\":\"Question\",\"name\":\"Why is a sentence split into several runs?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Because formatting changes within it. Each run holds characters with one consistent format, so a bold word in the middle of a sentence starts a new run.\"}},\n{\"@type\":\"Question\",\"name\":\"Can I read a DOCX without Microsoft Word?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Yes. Because the format is open XML, tools like the DOCX text extractor read it directly in the browser.\"}}\n]}<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>A .docx file is a ZIP archive of XML parts. The actual text lives in word\/document.xml, stored as a tree of paragraphs and &#8220;runs&#8221; of characters; images sit in word\/media\/, and formatting is kept\u2026<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"slim_seo":{"title":"How DOCX Files Store Their Content (It's a ZIP of XML)","description":"A .docx file is a ZIP of XML. The text lives in word\/document.xml as runs of characters, images sit in word\/media, and styles live separately. Here's the structure \u2014 and how to extract the text."},"footnotes":""},"categories":[3],"tags":[],"class_list":["post-56","post","type-post","status-publish","format-standard","hentry","category-guides"],"_links":{"self":[{"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/posts\/56","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/comments?post=56"}],"version-history":[{"count":1,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/posts\/56\/revisions"}],"predecessor-version":[{"id":77,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/posts\/56\/revisions\/77"}],"wp:attachment":[{"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/media?parent=56"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/categories?post=56"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/tags?post=56"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}