Which tool do you need?
| If you… | Use |
|---|---|
| You want the text out of a Word document | DOCX Text Extractor |
| You want to extract tables from a Word document | Word Table Extractor |
| You want the text and speaker notes from a PowerPoint | PowerPoint Text Extractor |
| You want the text of an EPUB ebook | EPUB Text Extractor |
| You want the images out of a Word, Excel or PowerPoint file | Office Image Extractor |
| Your document came out of Google Docs | Google Docs Image Extractor |
| You want to see who wrote it and when | Office Metadata Extractor |
These tools read the pre-2007 formats' successors only. A .doc, .xls or .ppt is a different, binary container and is not read here. To list everything inside any of these files raw, use the ZIP extractor.
Every modern Office file is a ZIP of XML
.docx, .pptx, .xlsx and .epub are all ZIP archives
holding XML. That single fact is why these tools are exact rather than approximate: the text is stored as
structured markup with headings, lists and reading order recorded explicitly, so it can be rebuilt rather
than flattened, and each embedded image is a complete file that comes out byte for byte. A PDF stores none
of this the same way, which is why PDF extraction is its own group.
The same structure is why metadata is worth checking before you send a document out. The author, the company on the licence and the name of whoever last saved the file live in two small XML parts inside every copy — read them with the Office metadata extractor before a file leaves your hands.
Every documents & office tool
- DOCX Text Extractor — Get the text out of a Word document as clean plain text or as Markdown that keeps headings, lists, tables and links intact.
- RTF Text Extractor — Get the plain text out of an RTF document — stripping all formatting control words while keeping paragraph breaks intact.
- ODT Text Extractor — Get the text out of a LibreOffice .odt document as plain text, keeping headings and paragraph order. Nothing uploaded.
- PowerPoint Text Extractor — Get the text off every slide and out of the speaker notes, in real presentation order, as plain text you can edit or search.
- EPUB Text Extractor — Pull an ebook's text out chapter by chapter in true reading order, with its title, author and publisher metadata.
- Office Image Extractor — Save every picture out of a Word, Excel or PowerPoint file at full original quality — not the downscaled copy Word gives you.
- Google Docs Image Extractor — Export your Doc as .docx and save every picture at full original quality — the sizes Google's own right-click download will not give you.
- Office Metadata Extractor — See the author, last-modified-by name, company, revision count and editing time hidden inside a Word, Excel or PowerPoint file.
- Word Table Extractor — Extract tables from Word documents (.docx) into clean CSV or JSON — cell spans expanded into a proper grid.
- Markdown Link Extractor — Extract all URLs, image sources, heading hierarchies, and code blocks from Markdown files into clean lists or CSV spreadsheets.
Frequently asked questions
What is the difference between the text and image extractors?
They read different parts of the same file. The text extractor rebuilds the words with their headings and lists; the image extractor pulls out the embedded pictures at full quality. Open the same .docx in either, depending on what you need.
Can these tools open old .doc or .ppt files?
No. Those pre-2007 formats are Compound File containers, not ZIPs, and store their content differently. These tools read .docx, .pptx, .xlsx and their relatives.
Do they upload my document?
No. Each file is unzipped and read inside your browser, which matters most for the metadata tool — the whole point there is to find what a document discloses without disclosing it.