What is inside an RTF file
RTF (Rich Text Format) was created by Microsoft in 1987 and became the standard interchange format
between word processors on different platforms. Underneath every RTF file is plain ASCII text mixed with
control words — backslash-prefixed keywords like \par (paragraph break), \b
(bold on), \fonttbl (font table definition) — and control symbols like \'e9
for a hex-encoded character.
The file is entirely text-based, which is why RTF is so portable: any text editor can open it. The problem is that those control words make up the majority of a typical RTF file — the actual reading content is buried inside them. This extractor removes the control layer and returns just the text.
How to extract text from an RTF file
- Open the file. Drop the .rtf file onto the box above, or click to browse. RTF files from any application are supported — Word, LibreOffice, WordPad, Pages, and older software.
- Read the extracted text. All body text is shown in the box below, with paragraph breaks preserved. Formatting characters, font tables, colour tables, and style sheets are removed.
- Copy or download. Copy the result to paste wherever you need it, or download it as a .txt file.
What comes out
- All body text in document order, with paragraph breaks preserved.
- Hex-encoded characters (
\'xx) decoded to real characters — so accented letters, smart quotes and non-ASCII text come through correctly. - Unicode escapes (
\usequences) decoded to the actual characters. - Word count shown after extraction.
Which files work
Any .rtf file produced by Word (any version), LibreOffice Write, WordPad, Apple Pages,
Google Docs (File → Download → RTF), or legacy word processors works here. RTF is the format that
lawyers, courts, and older enterprise systems exchange because it is universally readable.
The legacy binary .doc format and the XML-based .docx format are not RTF
— for those, use the DOCX text extractor. For OpenDocument
.odt files from LibreOffice, use the ODT text extractor.
Why the file is never uploaded
RTF is a plain-text format, so the extractor reads the file using a standard text decoder in your browser. No network request is made, no server is involved, and the file is gone when you close or refresh the tab.
RTF files are commonly generated from legal documents, medical records, and correspondence — exactly the type of content you should not hand to a server you cannot audit. Everything here runs locally.
What is not extracted
- Formatting — bold, italic, font sizes and colours are stripped by design. Only the reading text survives.
- Embedded images — RTF can embed images as hex or binary data; they are skipped.
- Headers and footers — these live in separate destinations and are not included.
- Tables — cell text is extracted but the grid layout is not reconstructed.
- Annotations and tracked changes — comment text is skipped.
Who uses RTF text extraction
- Legal and compliance teams — courts and law firms exchange RTF because it is application-neutral; getting the text into a case management system often requires stripping the markup.
- Writers and editors — manuscripts from older word processors arrive as RTF; extracting the text is the first step before importing into a modern editor.
- Developers — logging, search indexing, and migration tools need plain text, not RTF control words.
- Anyone without Word — reading an RTF file on a machine without a word processor.
RTF compared with DOCX
RTF and DOCX both came from Microsoft but are structurally opposite. RTF is a plain-text format where every byte is printable ASCII — you can open it in a text editor and read the control words alongside the text. DOCX is a ZIP archive of XML parts — binary at the container level, though the XML inside is readable.
RTF is older (1987 vs 2007), more portable across old software, and still the required exchange
format in many legal systems. DOCX has richer formatting, better table and image support, and is the
default in modern Word. For DOCX files use the DOCX text extractor.
For OpenDocument .odt files from LibreOffice use the
ODT text extractor.
Frequently asked questions
Can I open an RTF file without Word?
Yes — drop it onto this page. The file is read in your browser, so no Word, no LibreOffice and no software install is needed.
Does it handle accented characters and non-English text?
Yes. RTF encodes non-ASCII characters as hex escape sequences like \'e9. The extractor decodes them back to real characters, so French, German, Spanish and most Western-European text comes through correctly. Unicode escapes (\u sequences) are also decoded.
Is my RTF file uploaded to a server?
No. RTF is plain text, so it is read directly by your browser with no network request.
Why is there no Markdown output option?
RTF encodes formatting with control words rather than a semantic structure, so reconstructing Markdown headings and lists reliably is not possible. Plain text is what the format can reliably produce.
Does it work with RTF from LibreOffice or Pages?
Yes. RTF is a cross-application standard — files saved as .rtf from LibreOffice Writer, Apple Pages, or WordPad all use the same format and extract correctly.
What about tables inside the RTF?
Table cell text is extracted in left-to-right, top-to-bottom order, but the grid layout is not reconstructed — the cells come out as a flat sequence of paragraphs.