{"id":44,"date":"2026-08-07T10:26:00","date_gmt":"2026-08-07T10:26:00","guid":{"rendered":"https:\/\/easyextract.online\/blog\/what-is-inside-an-xlsx-file\/"},"modified":"2026-08-26T08:14:08","modified_gmt":"2026-08-26T08:14:08","slug":"what-is-inside-an-xlsx-file","status":"publish","type":"post","link":"https:\/\/easyextract.online\/blog\/what-is-inside-an-xlsx-file\/","title":{"rendered":"What&#8217;s Inside an XLSX File?"},"content":{"rendered":"<p><strong>An <code>.xlsx<\/code> file is not one binary blob \u2014 it is a ZIP archive containing a set of XML files. Rename a copy from <code>.xlsx<\/code> to <code>.zip<\/code>, open it, and you will see folders of plain-text XML that together describe the whole workbook.<\/strong> This structure \u2014 the Office Open XML format \u2014 is why <code>.xlsx<\/code> files can be read without Excel, and why extracting data from them is reliable. Here is what is inside.<\/p>\n<p>This guide walks through the parts of an XLSX, explains how cell values are actually stored, and shows how to <a href=\"https:\/\/easyextract.online\/xlsx-data-extractor\/\">extract data from Excel XLSX<\/a> in your browser without opening Excel at all.<\/p>\n<h2>An XLSX is a ZIP \u2014 try it yourself<\/h2>\n<p>Make a copy of any <code>.xlsx<\/code>, change its extension to <code>.zip<\/code>, and open it with your normal unzip tool. Instead of one file you will find a small tree of folders and XML documents. This packaging is called the <strong>Open Packaging Convention (OPC)<\/strong>, and the XML dialect inside is <strong>Office Open XML (OOXML)<\/strong>, standardised as ECMA-376. Word (<code>.docx<\/code>) and PowerPoint (<code>.pptx<\/code>) use the same idea.<\/p>\n<h2>The main parts and what they do<\/h2>\n<ul>\n<li><strong><code>[Content_Types].xml<\/code><\/strong> \u2014 a manifest listing the type of every part in the package. The reader consults it first.<\/li>\n<li><strong><code>_rels\/<\/code> and <code>xl\/_rels\/<\/code><\/strong> \u2014 relationship files that link the parts together (which worksheet is which, where the styles live).<\/li>\n<li><strong><code>xl\/workbook.xml<\/code><\/strong> \u2014 the workbook itself: the list of sheets, their names and order, defined names and settings.<\/li>\n<li><strong><code>xl\/worksheets\/sheet1.xml<\/code>, <code>sheet2.xml<\/code>\u2026<\/strong> \u2014 one file per worksheet, holding the rows, cells and their values.<\/li>\n<li><strong><code>xl\/sharedStrings.xml<\/code><\/strong> \u2014 a single table of every distinct piece of text in the workbook (explained below).<\/li>\n<li><strong><code>xl\/styles.xml<\/code><\/strong> \u2014 number formats, fonts, fills and borders, referenced by the cells rather than repeated in each one.<\/li>\n<li><strong><code>docProps\/core.xml<\/code> and <code>app.xml<\/code><\/strong> \u2014 document metadata: author, created and modified dates, the application that wrote the file.<\/li>\n<\/ul>\n<h2>How a cell value is actually stored<\/h2>\n<p>This is the part that surprises people. Open a worksheet XML and a cell looks like this:<\/p>\n<pre><code>&lt;c r=\"A1\" t=\"s\"&gt;&lt;v&gt;0&lt;\/v&gt;&lt;\/c&gt;<\/code><\/pre>\n<p>The <code>r=\"A1\"<\/code> is the cell reference. The <code>t=\"s\"<\/code> means the value is a <strong>shared string<\/strong>, and <code>&lt;v&gt;0&lt;\/v&gt;<\/code> is not the number zero \u2014 it is an <em>index<\/em> into <code>sharedStrings.xml<\/code>. To get the real text, the reader looks up entry <code>0<\/code> in the shared-strings table. Numbers are stored inline (<code>t<\/code> is absent and <code>&lt;v&gt;<\/code> holds the number), while text is de-duplicated into the shared table so a value repeated a thousand times is stored once.<\/p>\n<p>This is why you cannot just read a worksheet file on its own and expect to see your data \u2014 a correct extractor has to resolve shared strings, apply the number formats from <code>styles.xml<\/code>, and understand the relationships. That is the work a good tool does for you.<\/p>\n<h2>Two consequences worth knowing<\/h2>\n<ul>\n<li><strong>Dates are numbers.<\/strong> Excel stores a date as a serial number (days since 1900) and shows it as a date using a style. The raw value in the XML is the number \u2014 which is why exported data sometimes shows <code>45000<\/code> instead of a date.<\/li>\n<li><strong>Formulas keep a cached result.<\/strong> A formula cell stores both the formula and the last calculated value. Data extraction reads the cached value, so you get results, not the formula text.<\/li>\n<\/ul>\n<h2>How to get the data out<\/h2>\n<p>Because the format is open and text-based, you do not need Excel to read an XLSX. The <a href=\"https:\/\/easyextract.online\/xlsx-data-extractor\/\">Excel data extractor<\/a> opens the file in your browser, resolves the shared strings and styles for you, lets you pick the sheet and the columns you want, and exports clean CSV \u2014 with the workbook never leaving your device. If your data is already comma-separated, the <a href=\"https:\/\/easyextract.online\/csv-column-extractor\/\">CSV column extractor<\/a> does the same column-picking for CSV files.<\/p>\n<h2>Frequently asked questions<\/h2>\n<p><strong>Is an XLSX file really a ZIP?<\/strong><br \/>\nYes. An <code>.xlsx<\/code> is a ZIP archive of XML parts (the Office Open XML format). Rename a copy to <code>.zip<\/code> and you can browse the worksheets, shared strings and styles inside.<\/p>\n<p><strong>What is sharedStrings.xml for?<\/strong><br \/>\nIt is a single table of every distinct text value in the workbook. Cells store an index into this table instead of repeating the text, which keeps the file smaller. A reader has to resolve those indexes to show the real data.<\/p>\n<p><strong>Why does my exported date look like a number?<\/strong><br \/>\nExcel stores dates as serial numbers and displays them as dates using a style. The raw value is the number, so extraction can surface it as, for example, <code>45000<\/code> unless the format is applied.<\/p>\n<p><strong>Can I read an XLSX without Excel?<\/strong><br \/>\nYes. Because the format is open XML, tools like the <a href=\"https:\/\/easyextract.online\/xlsx-data-extractor\/\">Excel data extractor<\/a> read it directly in the browser and let you extract data from Excel XLSX without opening Excel.<\/p>\n<p><strong>What is the difference between XLSX and XLS?<\/strong><br \/>\n<code>.xlsx<\/code> is the modern ZIP-of-XML format. The older <code>.xls<\/code> is a single binary format and is unrelated; it has to be re-saved as <code>.xlsx<\/code> to use most modern XML-based tools.<\/p>\n<h2>Related reading<\/h2>\n<ul>\n<li><a href=\"https:\/\/easyextract.online\/blog\/csv-vs-xlsx\/\">CSV vs XLSX: what&#8217;s the difference and when to use each<\/a><\/li>\n<li><a href=\"https:\/\/easyextract.online\/blog\/how-to-extract-data-from-excel\/\">How to extract data from Excel<\/a><\/li>\n<\/ul>\n<p><em>Last updated: 16 August 2026.<\/em><\/p>\n<p><script type=\"application\/ld+json\">\n{\"@context\":\"https:\/\/schema.org\",\"@type\":\"FAQPage\",\"mainEntity\":[\n{\"@type\":\"Question\",\"name\":\"Is an XLSX file really a ZIP?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Yes. An .xlsx is a ZIP archive of XML parts (the Office Open XML format). Rename a copy to .zip and you can browse the worksheets, shared strings and styles inside.\"}},\n{\"@type\":\"Question\",\"name\":\"What is sharedStrings.xml for?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"It is a single table of every distinct text value in the workbook. Cells store an index into this table instead of repeating the text, which keeps the file smaller. A reader has to resolve those indexes to show the real data.\"}},\n{\"@type\":\"Question\",\"name\":\"Why does my exported date look like a number?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Excel stores dates as serial numbers and displays them as dates using a style. The raw value is the number, so extraction can surface it as a number like 45000 unless the format is applied.\"}},\n{\"@type\":\"Question\",\"name\":\"Can I read an XLSX without Excel?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"Yes. Because the format is open XML, tools like the Excel data extractor read it directly in the browser and let you extract data from Excel XLSX without opening Excel.\"}},\n{\"@type\":\"Question\",\"name\":\"What is the difference between XLSX and XLS?\",\"acceptedAnswer\":{\"@type\":\"Answer\",\"text\":\"XLSX is the modern ZIP-of-XML format. The older XLS is a single binary format and is unrelated; it has to be re-saved as XLSX to use most modern XML-based tools.\"}}\n]}<\/script><\/p>\n","protected":false},"excerpt":{"rendered":"<p>An .xlsx file is not one binary blob \u2014 it is a ZIP archive containing a set of XML files. Rename a copy from .xlsx to .zip, open it, and you will see folders\u2026<\/p>\n","protected":false},"author":2,"featured_media":0,"comment_status":"open","ping_status":"open","sticky":false,"template":"","format":"standard","meta":{"slim_seo":{"title":"What's Inside an XLSX File? A Look at the ZIP of XML Underneath","description":"An .xlsx file is a ZIP archive of XML parts. Rename it to .zip and look inside: worksheets, shared strings, styles and workbook. Here's what each part does, and how to extract data from Excel XLSX."},"footnotes":""},"categories":[3],"tags":[],"class_list":["post-44","post","type-post","status-publish","format-standard","hentry","category-guides"],"_links":{"self":[{"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/posts\/44","targetHints":{"allow":["GET"]}}],"collection":[{"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/posts"}],"about":[{"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/types\/post"}],"author":[{"embeddable":true,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/users\/2"}],"replies":[{"embeddable":true,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/comments?post=44"}],"version-history":[{"count":1,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/posts\/44\/revisions"}],"predecessor-version":[{"id":65,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/posts\/44\/revisions\/65"}],"wp:attachment":[{"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/media?parent=44"}],"wp:term":[{"taxonomy":"category","embeddable":true,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/categories?post=44"},{"taxonomy":"post_tag","embeddable":true,"href":"https:\/\/easyextract.online\/blog\/wp-json\/wp\/v2\/tags?post=44"}],"curies":[{"name":"wp","href":"https:\/\/api.w.org\/{rel}","templated":true}]}}