Documents & Office File Extraction

Office documents keep far more than the words on the page: images, speaker notes, revision history and the name of whoever last saved the file are all in there. These tools open a Word, PowerPoint, Excel or EPUB file and pull out exactly the part you want — the text, the pictures, or the metadata — without uploading the document anywhere.

Which tool do you need?

If you…Use
You want the text out of a Word documentDOCX Text Extractor
You want to extract tables from a Word documentWord Table Extractor
You want the text and speaker notes from a PowerPointPowerPoint Text Extractor
You want the text of an EPUB ebookEPUB Text Extractor
You want the images out of a Word, Excel or PowerPoint fileOffice Image Extractor
Your document came out of Google DocsGoogle Docs Image Extractor
You want to see who wrote it and whenOffice Metadata Extractor

These tools read the pre-2007 formats' successors only. A .doc, .xls or .ppt is a different, binary container and is not read here. To list everything inside any of these files raw, use the ZIP extractor.

Every modern Office file is a ZIP of XML

.docx, .pptx, .xlsx and .epub are all ZIP archives holding XML. That single fact is why these tools are exact rather than approximate: the text is stored as structured markup with headings, lists and reading order recorded explicitly, so it can be rebuilt rather than flattened, and each embedded image is a complete file that comes out byte for byte. A PDF stores none of this the same way, which is why PDF extraction is its own group.

The same structure is why metadata is worth checking before you send a document out. The author, the company on the licence and the name of whoever last saved the file live in two small XML parts inside every copy — read them with the Office metadata extractor before a file leaves your hands.

Every documents & office tool

Frequently asked questions

What is the difference between the text and image extractors?

They read different parts of the same file. The text extractor rebuilds the words with their headings and lists; the image extractor pulls out the embedded pictures at full quality. Open the same .docx in either, depending on what you need.

Can these tools open old .doc or .ppt files?

No. Those pre-2007 formats are Compound File containers, not ZIPs, and store their content differently. These tools read .docx, .pptx, .xlsx and their relatives.

Do they upload my document?

No. Each file is unzipped and read inside your browser, which matters most for the metadata tool — the whole point there is to find what a document discloses without disclosing it.

Other extraction groups