How Word documents store tables internally
A Word document (.docx) is an Open Packaging Convention (OPC) ZIP archive containing XML parts.
Tables are declared within word/document.xml using the <w:tbl> element, which contains
table rows (<w:tr>) and table cells (<w:tc>).
Cell merging in Word relies on two attributes: w:gridSpan for horizontal column spans and
w:vMerge for vertical row merges. Naive text copiers flatten these elements into continuous text strings,
losing all tabular structure. EasyExtract parses the OpenXML DOM tree, calculates exact grid coordinates for every cell,
and outputs aligned CSV representations.
For more details on Word internals, read how to extract images from Word or learn how text is structured with the DOCX text extractor.
How to extract tables from a Word document
- Select or drop a DOCX file. Drop your .docx file onto the box above or click to browse. The browser unzips the document container in memory without transmitting any data.
- Inspect discovered tables. The parser inspects word/document.xml, finds every w:tbl element, and displays a selector with row and column counts for each table.
- Review the grid. Grid spans (gridSpan) and merged cells (vMerge) are automatically expanded into a rectangular matrix to prevent column shifting when imported into spreadsheets.
- Export table data. Download the selected table as a CSV file, export all document tables into a single CSV, or copy the table as tab-separated values (TSV) to paste into Excel or Google Sheets.
What gets extracted
- All embedded tables, listed sequentially with row and column dimensions.
- Cell contents, with paragraph text and run breaks preserved as clean text.
- Expanded cell spans (column spans and vertical merges), ensuring every row has identical column counts.
- RFC 4180-compliant CSV output, properly quoting multi-line cells and commas.
- Tab-delimited clipboard output ready for instant spreadsheet pasting.
Supported formats and limits
This tool supports modern .docx files created by Microsoft Word, Google Docs, LibreOffice Writer,
and Pages. Password-protected files or legacy binary .doc files are not supported.
Because parsing is executed entirely via client-side JavaScript, file size is limited only by your device's available memory.
Privacy and security
Your document is parsed entirely inside your browser using client-side JavaScript and web decompression APIs. No file data, text, or table content is sent to any server.
Known limitations
Embedded images inside table cells are skipped (use our DOCX image extractor for media). Table styling, cell border colors, and background shading are omitted in favor of clean data output.
Common use cases
- Data Migration — Extracting pricing tables or financial matrices from Word reports into databases.
- Financial Analysis — Converting annual report tables from DOCX to Excel CSV format.
- Academic Research — Gathering survey result tables embedded in manuscripts.
Word Table Extractor vs HTML Table Extractor
While the HTML table extractor parses web markup (<table>),
this tool parses OpenXML markup inside Word archives. If you need to extract tables from spreadsheets, use the
Excel data extractor.
Frequently asked questions
How do I extract a table from Word to CSV?
Drop your .docx file onto the box, pick the table from the dropdown, and click Download CSV.
Does it support merged cells in Word tables?
Yes. Column spans (gridSpan) and row merges (vMerge) are expanded so the CSV columns align accurately.
Is my Word document uploaded to a server?
No. Parsing happens 100% in your browser. Zero document data leaves your device.
Can I extract all tables at once?
Yes. Clicking 'Download all tables' exports every table from the document into a single formatted CSV file.
Why are legacy .doc files not supported?
.doc is a legacy OLE2 binary format. Convert your file to .docx in Word or LibreOffice first.