What a regex extractor does
A regular expression describes a shape that text might take, and the engine returns every stretch of
your text fitting that shape. \d{4}-\d{2}-\d{2} describes four digits, a hyphen, two
digits, a hyphen, two digits — so it finds every ISO date in a document, regardless of what surrounds
them.
The difference from a fixed-purpose extractor is control. An email extractor matches one pattern someone chose. Here you choose it, which means you can pull out order references in your own company's format, log entries with a particular status code, or any value that has a shape but no standard.
How to extract text with a regular expression
- Paste your text. Any plain text works — a log file, an export, a page of HTML, a CSV, a document you have already converted to text.
- Write or choose a pattern. Type a regular expression, or pick a preset for emails, URLs, IP addresses, dates, hex colours, UUIDs and more, then adjust it.
- Choose what to return. Whole match returns everything the pattern matched. If your pattern has brackets, you can return just that capture group instead — the part inside the brackets.
- Extract and export. Results appear as a list, optionally deduplicated and sorted. Copy them or download as .txt or .csv.
What comes out
- Every match, in the order it appears, or sorted alphabetically.
- Capture groups — return just the part inside brackets rather than the whole match.
With
href="([^"]+)", the whole match includeshref="; group 1 is the URL alone. - Deduplicated results, so a value appearing 400 times is returned once.
- A match count, which doubles as a quick way to test a pattern.
- TXT and CSV export, with CSV properly quoted so matches containing commas survive.
Supported syntax and flags
The pattern engine follows the ECMAScript Regular Expression specification:
character classes, quantifiers, anchors, alternation, lookahead and lookbehind, named and numbered groups, and Unicode escapes. Three flags are
exposed as checkboxes — ignore case, multiline (so ^ and $ match at each line
break) and dotall (so . also matches newlines). Global matching is always on, because
extraction means finding every occurrence.
Fourteen presets cover the common cases: email addresses, URLs, IPv4, hex colours, ISO dates, times, currency amounts, hashtags, mentions, quoted strings, HTML tag names, UUIDs and numbers. Each loads into the pattern box so you can adapt it rather than start from nothing. Matching stops at 200,000 results, which is reported rather than silently truncated.
Why your text and pattern stay on your device
Matching runs inside your browser, in a Web Worker. Neither the text nor the pattern is transmitted. Online regex tools that evaluate server-side receive both — and the text people test regexes against is frequently a real log file or a real export.
The worker also serves a second purpose, which is safety rather than privacy: see below.
Runaway patterns, and what the time limit protects you from
Some patterns take effectively forever. (a+)+$ against a long line that does not match
forces the engine to try an exponential number of combinations — catastrophic backtracking. JavaScript
offers no way to interrupt a regex once it starts, so a page running one directly becomes permanently
unresponsive and has to be closed.
This tool runs every pattern in a Web Worker, a separate thread, and terminates that worker after three seconds. A runaway pattern costs you a message explaining what happened, not the tab and everything you had typed into it.
Two other boundaries worth stating: regular expressions cannot reliably parse nested structures such as HTML or JSON, because nesting has no fixed depth — use the document extractors or a real parser for those. And replacement is not offered here; this tool extracts, it does not rewrite.
Who uses a regex extractor
- Developers — pulling ids, timestamps or error codes out of a log without writing a script.
- Data teams — extracting a field from a semi-structured export that no dedicated tool understands.
- SEO and content work — lifting every URL, heading or attribute value out of an HTML dump.
- Anyone testing a pattern — checking a regex against real text before putting it into code.
- Support engineers — isolating the lines of a customer log that match a signature.
When to use a preset tool instead
Use a dedicated extractor when one exists. The email and URL extractor already handles the awkward edges of address matching, and the phone number extractor does far more than a pattern can — it validates digit counts and converts to international format, which no regex will do for you.
Reach for this tool when your values have a shape but no standard: internal reference codes, a log format specific to your systems, a field inside a proprietary export. That is exactly the gap a general-purpose extractor fills.
To master parentheses and numbered groups for isolating sub-strings from matched patterns, read regex capture groups for data extraction.
Regex engine: ECMAScript standards and known limits
The regex engine is the one built into your browser, following the ECMAScript
specification. Supported features include character classes, non-capturing groups, lookahead,
lookbehind (ES2018+), named capture groups (ES2018+) and Unicode property escapes with the
/u flag (ES2018+). In browsers that shipped before ES2018, lookbehind assertions
are absent without error.
Three edge cases to know before running a pattern: (a) catastrophic backtracking —
patterns with nested quantifiers such as (a+)+ can take exponential time on
long inputs; the extractor imposes no timeout, so a poorly written pattern on a large paste
will freeze the tab; (b) the /g flag is added automatically — writing your
pattern without it returns the same result as writing it with it; (c) without the
/m flag, ^ and $ match the very start and end of the
full input, not individual lines — add /m explicitly to match line
boundaries.
Frequently asked questions
How do I extract text using a regular expression online?
Paste your text, type a pattern or choose a preset, and click Extract. Every match is listed and can be copied or downloaded.
Is my text sent to a server?
No. Both the text and the pattern stay in your browser — matching runs in a Web Worker on your own machine.
What happens if my pattern is very slow?
It is stopped after three seconds and you get an explanation. Because the pattern runs in a separate worker thread, the page itself never freezes and your text is still there.
How do I return only part of a match?
Put brackets around the part you want and select that capture group in the Return dropdown. With href="([^"]+)", group 1 gives you the URL without the surrounding markup.
Which regex syntax does it use?
JavaScript's, which is close to PCRE for everyday patterns. Lookahead, lookbehind and named groups all work. Global matching is always enabled.
Why does my ^ anchor not match every line?
By default ^ and $ match only the very start and end of the text. Tick “^ $ match each line” to make them match at every line break.
Can I use it to replace text as well?
No, this tool extracts only. Use a text editor with regex find-and-replace for rewriting.