How a .pptx stores its text
A .pptx is a ZIP of XML parts, one per slide, plus a presentation part that lists slide
ids. Those ids are resolved through a relationship file to actual slide paths, and that list — the
sldIdLst — is the real running order. Slide filenames are assigned when slides are created,
so a deck that has been reordered has filenames in a completely different sequence from the presentation.
Text itself sits in DrawingML: each shape holds paragraphs, each paragraph holds runs of characters. Speaker notes live in separate notesSlide parts, linked from each individual slide rather than numbered in parallel, which is why matching notes to slides by number is unreliable.
How to extract text from a PowerPoint file
- Open the presentation. Drop the .pptx onto the box above, or click to browse. Slides are read locally in presentation order, not filename order.
- Choose what to include. Label each slide adds a numbered separator with the slide's title. Include speaker notes appends the notes beneath each slide's text.
- Review the output. Slide text appears shape by shape, in the order PowerPoint stores the shapes, with each paragraph on its own line.
- Copy or download. Copy the whole deck to the clipboard, or download it as a .txt file.
What comes out
- All slide text — titles, body placeholders, text boxes and text inside tables and grouped shapes.
- Correct slide order, resolved through the presentation's relationships rather than guessed from filenames.
- Speaker notes for each slide, resolved through that slide's own relationships.
- Slide labels using each slide's first line as its title, so the output reads as an outline.
- Counts — slides, words, how many slides carry notes, and how many contain no text at all.
Slides that are entirely images are reported rather than silently skipped, so a suspiciously short result explains itself.
Which presentations work
.pptx and .ppsx files from PowerPoint 2007 onward work, along with decks
exported by Google Slides, Keynote and LibreOffice Impress. Only slide XML is parsed, so a deck full of
high-resolution images opens as fast as a text-only one.
The legacy .ppt format is binary and unrelated; it is detected and reported rather than
misread. If the relationship list is missing or damaged, the tool falls back to slide filenames sorted
numerically — so slide 10 still lands after slide 9 rather than after slide 1.
Why the deck is never uploaded
The archive is read from your disk by the browser and decompressed locally. No server takes part, so the presentation is never transmitted and nothing survives closing the tab.
Decks are routinely more sensitive than they look: board presentations, pricing, unreleased roadmaps, and speaker notes that were never meant to leave the room. Notes in particular are the part people forget is in the file at all.
What is not included
Plain text cannot carry a deck's visual content, so five things are absent:
- Images, charts and SmartArt. Text labels inside SmartArt are extracted; the graphics are not.
- Slide masters and layouts. Placeholder text that lives on the layout rather than the slide is not part of the slide's own content.
- Comments and revision history.
- Animations, transitions and timings.
- Reading order within a slide. Shapes come out in the order they are stored, which is creation order — usually but not always top-to-bottom on screen.
Text inside embedded objects, such as a linked spreadsheet, also stays inside that object.
Who extracts text from presentations
- Turning a deck into a document — recovering an outline to write up as prose.
- Translation — getting every string out for a translator, including notes.
- Search and audit — making a deck's content greppable, or checking which slides mention a term.
- Accessibility — producing a text transcript of a presentation.
- Feeding a deck to an AI tool — most models cannot read .pptx directly.
- Recovering notes — pulling the speaker notes out of a deck someone else built.
Slide text compared with document text
A deck is a set of fragments; a document is continuous prose. Extraction reflects that — slide output reads as an outline of short lines, not paragraphs, and no amount of processing will turn one into the other.
For a Word document, the DOCX text extractor keeps headings, lists and tables, and can output Markdown. For a PDF export of a deck, use the PDF text extractor. To pull the pictures out of the deck instead, use the Office image extractor, which reads PowerPoint files too.
To understand the hidden revision histories, author details, and comments stored in Office files, read our guide on hidden metadata in Office files.
PPTX XML structure and extraction edge cases
A PPTX is an OOXML package per ISO/IEC 29500. Slides are stored as
ppt/slides/slide*.xml. Text runs live inside shape text bodies:
<p:sp><p:txBody><a:p><a:r><a:t>. Slide reading
order is determined by the <p:sldIdLst> list in
ppt/presentation.xml, not by filename.
For a guide to what is stored inside Office Open XML files, see the article on hidden metadata in Office files. To extract images from a PowerPoint file instead of text, use the Office image extractor.
Three edge cases specific to PowerPoint: (a) placeholder text — "Click to add title" — is stored in the slide layout and master XMLs, not the slide file itself; a slide that was never edited shows no text for those placeholders; (b) animation order is not reading order — a slide may visually reveal bullet 3 before bullet 1, but extraction returns elements in XML document order (the author's editing order); (c) SmartArt is stored as DrawingML with a text node at each shape level — the extractor returns each text node, which often produces more lines than the visible SmartArt diagram suggests.
Frequently asked questions
How do I extract text from a PowerPoint without PowerPoint?
Drop the .pptx onto this page. It is read in your browser, so no Office licence and no install are needed.
Does it include speaker notes?
Yes, and they are matched to slides through each slide's own relationships rather than by number, which is why they line up correctly even in reordered decks.
Is my presentation uploaded to a server?
No. The .pptx is read from your disk and unzipped by your browser. It is never transmitted.
Are the slides in the right order?
Yes. Order comes from the presentation's slide id list, not from filenames. A deck whose slides have been reordered still extracts in presentation order.
Why is one of my slides empty in the output?
That slide contains no text — it is an image, a chart or a diagram. The tool reports how many slides are image-only so a short result is explained.
Can it extract the images too?
Not this tool. Use the Office image extractor, which reads the media folder inside PowerPoint, Word and Excel files.
Why does it reject my .ppt file?
The old .ppt format is binary and completely different from .pptx. Open it in PowerPoint or LibreOffice and save it as .pptx first.
Does it work on Mac, Linux or a phone?
Yes. Everything runs in the browser, so no operating system or Office install is required.