Supported Extraction File Formats
Last updated: 2 August 2026
EasyExtract's online extraction tools support the formats listed on this page. Every tool runs entirely in your browser — the file is read from your device into browser memory and parsed there. Nothing is transmitted to a server. Size limits are therefore set by your device's available RAM, not by any server-side upload cap.
The tables below list each accepted extension, its MIME type, the practical tested size range, and what the tool produces. If your format is not listed, let us know.
All four PDF tools share the same parser (PDF.js, loaded from EasyExtract's own domain). They accept any conforming PDF — text-layer, scanned, or mixed — though scanned-only files will show no text in the text and table tools. Password-protected PDFs must be unlocked in a PDF reader before use.
| Format | Extension(s) | MIME type | Practical size | Tool | Output |
|---|---|---|---|---|---|
| PDF (text layer) | .pdf |
application/pdf |
Tested to 200 MB; larger files depend on RAM | PDF Text Extractor | Plain text (.txt) |
| PDF (tables) | .pdf |
application/pdf |
Tested to 200 MB | PDF Table Extractor | CSV or TSV |
| PDF (metadata) | .pdf |
application/pdf |
Reads only the metadata block — any size | PDF Metadata Extractor | Key–value table |
| PDF (embedded images) | .pdf |
application/pdf |
Tested to 200 MB | PDF Image Extractor | PNG or JPEG images (download individually or as ZIP) |
Documents & Office
Office Open XML files (.docx, .pptx, .xlsx) are ZIP archives containing XML. The tools unzip and parse the XML in your browser using file extension vs MIME type detection to confirm the format before processing. Legacy binary Office formats (.doc, .xls, .ppt) are not supported.
| Format | Extension(s) | MIME type | Practical size | Tool | Output |
|---|---|---|---|---|---|
| Word Document | .docx .docm .dotx |
application/vnd.openxmlformats-officedocument.wordprocessingml.document |
Tested to 50 MB | DOCX Text Extractor | Plain text (.txt) |
| PowerPoint Presentation | .pptx .ppsx |
application/vnd.openxmlformats-officedocument.presentationml.presentation |
Tested to 50 MB | PowerPoint Text Extractor | Plain text (.txt), one slide per section |
| EPUB eBook | .epub |
application/epub+zip |
Tested to 30 MB | EPUB Text Extractor | Plain text (.txt) |
| Office images (DOCX, XLSX, PPTX, ODT, ODS, ODP, ODG) | .docx .docm .dotx .xlsx .xlsm .xltx .pptx .pptm .ppsx .odt .ods .odp .odg |
Various (all ZIP-based) | Tested to 100 MB | Office Image Extractor | PNG or JPEG images (download individually or as ZIP) |
| Office metadata (DOCX, XLSX, PPTX) | .docx .xlsx .pptx .docm .xlsm .pptm |
Various OOXML MIME types | Reads only docProps/ — any size |
Office Metadata Extractor | Key–value table (author, dates, revision count, etc.) |
Spreadsheets & structured data
Structured data tools accept common data-interchange formats. The HTML tools accept pasted markup rather than a file upload, so there is no file-size limit for those — only browser rendering limits apply.
| Format | Extension(s) | MIME type | Practical size | Tool | Output |
|---|---|---|---|---|---|
| Excel Workbook | .xlsx .xlsm |
application/vnd.openxmlformats-officedocument.spreadsheetml.sheet |
Tested to 30 MB / ~500k rows | Excel Data Extractor | CSV (per sheet) |
| CSV / TSV | .csv .tsv .txt |
text/csv |
Tested to 50 MB | CSV Column Extractor | CSV (selected columns) |
| JSON / GeoJSON | .json .geojson .txt |
application/json |
Tested to 20 MB | JSON Field Extractor | Extracted values as text or JSON |
| HTML (tables) | Paste or upload .html |
text/html |
Paste any size | HTML Table Extractor | CSV (one per <table>) |
| HTML (text) | Paste or upload .html |
text/html |
Paste any size | HTML Text Extractor | Plain text (tags stripped) |
Archives & executables
Archive tools list contents and allow you to save individual entries. Only the entries you choose to save are fully decompressed in memory. The ZIP tool can open any ZIP-compatible container — including EPUB, DOCX, JAR, and APK — because those formats are ZIP archives with different extensions.
| Format | Extension(s) | MIME type | Practical size | Tool | Output |
|---|---|---|---|---|---|
| ZIP (and ZIP-compatible) | .zip .jar .apk .epub .docx .xlsx .pptx .odt |
application/zip |
Up to 4 GB (ZIP64 extended offsets reported, not read) | ZIP Extractor | File listing; individual files on demand |
| TAR / TAR.GZ | .tar .gz .tgz .taz |
application/x-tar application/gzip |
Limited by RAM; tested to 500 MB | TAR Extractor | File listing; individual files on demand |
| Windows executable / DLL | .exe .dll .sys .ocx .scr .cpl .node |
application/vnd.microsoft.portable-executable |
Tested to 200 MB | EXE Extractor | PE headers, section table, embedded strings |
| Windows icon source (EXE / DLL) | .exe .dll .ico .cur .icl .ocx .cpl .scr .mun |
image/x-icon or PE MIME |
Tested to 200 MB | Icon Extractor | ICO or PNG files (each size variant) |
| Windows Installer | .msi .msp .msm .mst |
application/x-msi |
Tested to 500 MB | MSI Extractor | File listing, property table, summary stream |
Images
Three separate tools process image files, each with a different goal. The EXIF tool reads only JPEG and TIFF, because those are the formats that carry EXIF metadata. The OCR and colour-palette tools accept any image the browser can decode.
| Format | Extension(s) | MIME type | Practical size | Tool | Output |
|---|---|---|---|---|---|
| JPEG / TIFF (EXIF metadata) | .jpg .jpeg .tif .tiff |
image/jpeg image/tiff |
Only the EXIF header is read — any file size | EXIF Data Extractor | EXIF field table (camera, lens, GPS, dates) |
| Any browser-renderable image (OCR) | .jpg .jpeg .png .gif .webp .bmp .avif |
image/* |
Tested to 10 MB; very large images slow Tesseract.js | Image to Text (OCR) | Plain text |
| Any browser-renderable image (colour palette) | .jpg .jpeg .png .gif .webp .bmp .avif |
image/* |
Tested to 10 MB | Colour Palette Extractor | Hex colour swatches + CSV |
Audio
| Format | Extension(s) | MIME type | Practical size | Tool | Output |
|---|---|---|---|---|---|
| MP3 | .mp3 |
audio/mpeg |
Only the ID3 tag block is read — any file size | MP3 Tag & Album Art Extractor | ID3 tag table; album art as downloadable image |
Email files
| Format | Extension(s) | MIME type | Practical size | Tool | Output |
|---|---|---|---|---|---|
| Outlook Message | .msg |
application/vnd.ms-outlook |
Tested to 20 MB | Outlook MSG Viewer | Subject, body, headers, sender, recipients |
| EML / MHTML | .eml .mht |
message/rfc822 |
Tested to 20 MB | EML Viewer | Subject, body, headers, sender, recipients |
Text, patterns & subtitles
Text-pattern tools extract specific data types (emails, phone numbers, IP addresses, URLs, numbers) from pasted text. They do not use a file picker. The subtitle tool accepts common caption formats as a file upload.
| Input type | Extension(s) | Practical limit | Tool | Output |
|---|---|---|---|---|
| Pasted text (email addresses) | n/a — paste only | Browser textarea limit (~millions of characters) | Email & URL Extractor | Unique email list (.txt) |
| Pasted text (phone numbers) | n/a — paste only | Browser textarea limit | Phone Number Extractor | Phone number list (.txt) |
| Pasted text (custom regex) | n/a — paste only | Browser textarea limit | Regex Extractor | Match list (.txt) |
| Pasted text (IP addresses) | n/a — paste only | Browser textarea limit | IP Address Extractor | IPv4 and IPv6 list (.txt) |
| Pasted text (URLs) | n/a — paste only | Browser textarea limit | URL Extractor | URL list (.txt) |
| Pasted text (numbers) | n/a — paste only | Browser textarea limit | Number Extractor | Number list (.txt) |
| Subtitle / caption file | .srt .vtt .sbv .sub |
Any size | SRT Subtitle Viewer | Cue list; plain-text transcript (.txt) |
Notes on size limits
Because all processing happens in your browser, there is no server-enforced upload cap. The limits shown above are the sizes tested during development. Files beyond those limits may work fine, or they may be slow or fail depending on how much RAM your device has available. If a large file causes the page to hang, reload the tab — closing the tab discards everything immediately with no data leaving your device.
Formats that only read a header block — EXIF, MP3 ID3, PDF metadata, and Office metadata — load almost instantly regardless of file size, because only the first few kilobytes are accessed.
Notes on MIME types and extensions
The browser file picker uses the accept attribute to suggest which files to show, but the actual format is confirmed by reading the file's magic bytes before processing begins. This means a file renamed with the wrong extension will be caught before any parsing attempt. For a full explanation of how file extensions differ from MIME types, see our guide on file extension vs MIME type.
Formats not supported
The following are the most commonly requested formats that are not currently supported:
- Legacy Office formats —
.doc,.xls,.ppt(binary compound document format; not supported by any current tool) - RAR archives — the RAR format requires a proprietary parser that is not available as a permissively licensed browser library
- 7-Zip archives — not currently supported
- Password-protected files — any format that requires a decryption key before parsing is not supported; unlock the file locally first
- Video files — metadata extraction from MP4, MKV, or MOV is not currently available
If a format you need is missing, send a request via the contact page.
Ready to extract? Every tool runs in your browser.
No account, no upload, no server. Your file stays on your device from start to finish.
Browse all online extraction tools →