Why reading an XLSX is not simply reading cells
An .xlsx is a ZIP of XML parts. Cell values are not all stored in the cells: any repeated
piece of text is written once into xl/sharedStrings.xml, and each cell that uses it stores an
index into that table. A cell containing "EMEA" may hold nothing but the number 7.
Extraction therefore has two steps — resolve the shared string table, then walk each sheet and place values by their cell reference. Cell references matter because sparse sheets skip empty cells entirely: a row with data in A and D contains two cells, not four, so anything that reads them in order shifts your columns left.
For a full breakdown of every part stored inside an .xlsx archive, read what is inside an XLSX file.
How to extract data from an Excel file
- Open the workbook. Drop the .xlsx onto the box above, or click to browse. Every sheet is read locally, with the shared string table resolved so text appears as text.
- Pick a sheet. The selector lists every sheet with its row count. Switching sheets re-renders the preview instantly — nothing is re-read.
- Check the grid. The preview shows the first 200 rows and reports the true row and column count, so you can confirm the data landed where you expect.
- Export. Download the current sheet as CSV, every sheet in one file, or copy the sheet as tab-separated values to paste straight into another spreadsheet.
What comes out
- Every sheet in the workbook, in workbook order, with its real name.
- All cell values — text resolved from the shared string table, numbers, booleans as TRUE/FALSE, and the cached results of formulas.
- Correct column alignment, because values are placed by cell reference rather than by the order they appear in the file.
- Properly quoted CSV, so values containing commas, quotes or line breaks survive the round trip into Excel or Google Sheets.
- Tab-separated output on the clipboard, which pastes into a spreadsheet as real columns.
Which workbooks work
.xlsx and .xlsm files from Excel 2007 onward work, as do workbooks exported by
Google Sheets, LibreOffice and Numbers. There is no row or sheet limit beyond your device's memory, and
only the XML is parsed, so charts and images cost nothing.
The legacy .xls format is binary and completely different; it is detected and reported
rather than misread. Password-protected workbooks are encrypted at the container level and must be
unlocked first.
Why the workbook is never uploaded
The archive is read from your disk by the browser and decompressed with the browser's own DecompressionStream. No server is involved, so the workbook is never transmitted and no copy remains once the tab closes.
Spreadsheets are the highest-risk file type to hand to an online converter, because they are where organisations keep customer lists, salary data, pricing and financial models. A single upload of the wrong workbook is a data incident. Reading it locally removes that possibility.
What is not converted
CSV is a grid of values, so five things a workbook holds do not survive:
- Formulas. The cached result is exported, not the formula. A workbook saved by a tool that did not cache results will show blanks in those cells.
- Date formatting. Excel stores dates as serial numbers and applies a display format,
so a date may export as
45367. Format the column as text before saving, or convert the serial afterwards. - Merged cells. The value sits in the top-left cell of the merge; the rest are empty.
- Formatting, colours and conditional rules. None of these are values.
- Charts, pivot tables and macros. Pivot caches are not expanded.
Hidden rows, columns and sheets are extracted, which is worth knowing before sharing the output — hidden does not mean absent.
Who extracts data from Excel files
- Developers and analysts — turning a supplied workbook into CSV a script or database can load.
- People without Excel — reading a workbook on a machine or phone with no Office licence.
- Data migration — moving sheets into a system that only accepts CSV.
- Auditing a file before trusting it — checking every sheet, including hidden ones, for content you did not expect.
- Feeding a spreadsheet to an AI tool — producing plain CSV a model can read.
Extracting from a workbook compared with a PDF table
A workbook already knows its own structure: rows, columns and sheets are explicit, so extraction is exact. A table inside a PDF has no structure at all — only text at coordinates — so the PDF table extractor has to infer columns from gaps between words, and sometimes gets it wrong. If your data exists as both, always take the workbook.
To read the text of a Word document instead, use the DOCX text extractor. To pull the images out of a workbook, use the Office image extractor. To see the raw XML parts inside the container, open it with the ZIP extractor.
For a full walkthrough of every method — copy-paste, CSV export, Power Query, formulas, and this tool — read how to extract data from Excel.
XLSX data storage and extraction edge cases
An XLSX workbook stores string values in a Shared String Table (SST) at
xl/sharedStrings.xml — a string cell holds an integer index into the SST rather
than the text itself. Numbers and dates are stored as IEEE 754 double-precision floats; a cell
displays as a date because of its number-format code, not because it holds a date type.
Three extraction edge cases: (a) merged cells write the value to the top-left cell only —
the extractor returns the value at the origin and leaves the merged range empty, so a merged
header spanning four columns appears once in the CSV output; (b) formula results are available
only if the workbook was saved with cached values — if a file was closed in manual-recalculation
mode, formula cells appear blank; (c) integers above 253 lose precision because they
are stored as IEEE 754 doubles — the last few digits may differ from what Excel displayed at
the time of creation, since JavaScript's Number type has the same 53-bit mantissa
limit.
Frequently asked questions
How do I convert XLSX to CSV without Excel?
Drop the workbook onto this page, choose a sheet and click Download this sheet as CSV. No Excel licence and no install are needed.
Can it extract all sheets at once?
Yes. Download all sheets writes one file with each sheet's CSV in order, separated by a labelled header line.
Is my spreadsheet uploaded to a server?
No. The workbook is read from your disk and unzipped by your browser. It is never transmitted.
Why do my dates look like numbers?
Excel stores a date as a serial number and applies a display format on top. CSV carries values, not formats, so the serial is exported. Format the column as text in Excel before saving if you need the readable date.
Do I get formulas or their results?
Their results. Excel caches the last calculated value of each formula and that is what is exported. The formula itself is not.
Are hidden sheets and rows included?
Yes — all of them. Hidden content is still data in the file, so check the output before sharing it.
Why does it reject my .xls file?
The old .xls format is binary and unrelated to .xlsx. Open it in Excel or LibreOffice and save it as .xlsx first.
Is there a row limit?
No. Parsing happens locally, so the only limit is your device's memory. The preview shows 200 rows; the CSV contains every row.