Guides

What’s Inside an XLSX File?

An .xlsx file is not one binary blob — it is a ZIP archive containing a set of XML files. Rename a copy from .xlsx to .zip, open it, and you will see folders of plain-text XML that together describe the whole workbook. This structure — the Office Open XML format — is why .xlsx files can be read without Excel, and why extracting data from them is reliable. Here is what is inside.

This guide walks through the parts of an XLSX, explains how cell values are actually stored, and shows how to extract data from Excel XLSX in your browser without opening Excel at all.

An XLSX is a ZIP — try it yourself

Make a copy of any .xlsx, change its extension to .zip, and open it with your normal unzip tool. Instead of one file you will find a small tree of folders and XML documents. This packaging is called the Open Packaging Convention (OPC), and the XML dialect inside is Office Open XML (OOXML), standardised as ECMA-376. Word (.docx) and PowerPoint (.pptx) use the same idea.

The main parts and what they do

  • [Content_Types].xml — a manifest listing the type of every part in the package. The reader consults it first.
  • _rels/ and xl/_rels/ — relationship files that link the parts together (which worksheet is which, where the styles live).
  • xl/workbook.xml — the workbook itself: the list of sheets, their names and order, defined names and settings.
  • xl/worksheets/sheet1.xml, sheet2.xml… — one file per worksheet, holding the rows, cells and their values.
  • xl/sharedStrings.xml — a single table of every distinct piece of text in the workbook (explained below).
  • xl/styles.xml — number formats, fonts, fills and borders, referenced by the cells rather than repeated in each one.
  • docProps/core.xml and app.xml — document metadata: author, created and modified dates, the application that wrote the file.

How a cell value is actually stored

This is the part that surprises people. Open a worksheet XML and a cell looks like this:

<c r="A1" t="s"><v>0</v></c>

The r="A1" is the cell reference. The t="s" means the value is a shared string, and <v>0</v> is not the number zero — it is an index into sharedStrings.xml. To get the real text, the reader looks up entry 0 in the shared-strings table. Numbers are stored inline (t is absent and <v> holds the number), while text is de-duplicated into the shared table so a value repeated a thousand times is stored once.

This is why you cannot just read a worksheet file on its own and expect to see your data — a correct extractor has to resolve shared strings, apply the number formats from styles.xml, and understand the relationships. That is the work a good tool does for you.

Two consequences worth knowing

  • Dates are numbers. Excel stores a date as a serial number (days since 1900) and shows it as a date using a style. The raw value in the XML is the number — which is why exported data sometimes shows 45000 instead of a date.
  • Formulas keep a cached result. A formula cell stores both the formula and the last calculated value. Data extraction reads the cached value, so you get results, not the formula text.

How to get the data out

Because the format is open and text-based, you do not need Excel to read an XLSX. The Excel data extractor opens the file in your browser, resolves the shared strings and styles for you, lets you pick the sheet and the columns you want, and exports clean CSV — with the workbook never leaving your device. If your data is already comma-separated, the CSV column extractor does the same column-picking for CSV files.

Frequently asked questions

Is an XLSX file really a ZIP?
Yes. An .xlsx is a ZIP archive of XML parts (the Office Open XML format). Rename a copy to .zip and you can browse the worksheets, shared strings and styles inside.

What is sharedStrings.xml for?
It is a single table of every distinct text value in the workbook. Cells store an index into this table instead of repeating the text, which keeps the file smaller. A reader has to resolve those indexes to show the real data.

Why does my exported date look like a number?
Excel stores dates as serial numbers and displays them as dates using a style. The raw value is the number, so extraction can surface it as, for example, 45000 unless the format is applied.

Can I read an XLSX without Excel?
Yes. Because the format is open XML, tools like the Excel data extractor read it directly in the browser and let you extract data from Excel XLSX without opening Excel.

What is the difference between XLSX and XLS?
.xlsx is the modern ZIP-of-XML format. The older .xls is a single binary format and is unrelated; it has to be re-saved as .xlsx to use most modern XML-based tools.

Last updated: 16 August 2026.

About Abrar

Abrar builds EasyExtract's free, browser-based extraction tools and writes these guides on getting data out of files — PDFs, spreadsheets, images, archives and Office documents. Every tool runs entirely in your browser, so nothing you open is ever uploaded.

Keep reading