What counts as a URL here
A URL is a web address, and in real text it appears in several shapes. The fully written form is
https://acme.com/pricing?ref=home. Just as common are the shorthand forms people actually
type: www.acme.com with no scheme, or a bare acme.com/blog with neither. This
tool recognises all three, because text copied from the wild is full of the shorthand ones and a tool that
only matched https:// would miss half the links on the page.
Each match is then parsed properly rather than by pattern alone. The address is resolved with the browser’s own URL parser, which is how the domain is pulled out reliably and how tracking parameters are identified and removed — string matching cannot tell a real query parameter from part of a path, but a parser can.
How to extract URLs from text
- Paste the text. Drop any text into the box above — an email, a page of HTML source, a chat log, a document. The links do not need to be formatted or on their own lines.
- Choose the view. Every URL lists each link in full. Unique domains rolls them up to one line per site, with a count of how many links point there — useful for auditing what a page links out to.
- Tune the result. Strip tracking removes utm and click-id parameters. Include links with no https:// catches bare addresses like acme.com/blog. Pages only drops image, CSS and script URLs.
- Copy or export. Copy list puts the current view on your clipboard. Save as CSV writes each URL with its domain, or each domain with its link count, ready for a spreadsheet.
What comes out
- Every link in the text, in full, with trailing sentence punctuation trimmed so a URL at the end of a sentence is not left with a stray full stop.
- A unique-domain rollup — one line per site, ordered by how many links point to it, which turns a wall of links into a short list of who a page actually references.
- Cleaned URLs, with
utm_*,fbclid,gclidand the other common tracking parameters stripped, so the list is the real destinations rather than campaign noise. - Bare addresses written without a scheme, caught alongside the fully written ones.
Exports keep both dimensions: the CSV of URLs carries each link’s domain in a second column, and the CSV of domains carries each domain’s link count — so the result is usable in a spreadsheet without re-parsing.
What is matched, and what is not
Matched: http and https links, protocol-relative
//host/path links, www. addresses, and — when the option is on —
bare domains with a valid public suffix such as acme.com or example.co.uk.
Tracking parameters are removed using the browser’s URL parser, so only genuine query parameters
are touched and the rest of the address is left exactly as written.
Not matched: email addresses, which are a different thing — to pull those out,
use the email and URL extractor. mailto:,
tel:, ftp: and javascript: links are also skipped, because the
job here is web addresses. The bare-domain pass is deliberately conservative: it requires a real
top-level domain, so a sentence like “version 2.0 is out” is not mistaken for a link.
Why the text is never uploaded
Everything happens in your browser. The text you paste is matched and parsed on your own device, and no server takes part, so nothing is transmitted.
That matters when the text is not public. Pulling the links out of a private email thread, an internal document or a page of someone else’s HTML source are all routine uses, and none of them should involve sending that text to a stranger’s server first. Here there is nothing to send.
What this tool does not do
Three boundaries worth stating:
- It does not fetch or check the links. It reads them out of the text you paste; it does not visit them, so it cannot tell you which are live or where a shortener points.
- It does not expand shortened URLs. A
bit.lylink comes out as thebit.lylink, because resolving it would mean fetching it from a server. - It does not read links out of a file. Paste the text in. To get the text out of a PDF or a document first, use the PDF text extractor or the DOCX text extractor, then bring the result here.
Why people extract URLs from text
- Auditing outbound links — pasting a page’s text or source to see every site it points to, grouped by domain.
- Cleaning a list — pulling the real destinations out of a marketing email where every link is wrapped in tracking parameters.
- Research and citations — collecting every source linked in an article into one list.
- Migration and QA — extracting the links from old content to check them before a site move.
- Security review — reading the links out of a suspicious message without clicking any of them, and seeing which domains they really go to.
The URL extractor compared with the email extractor
The email and URL extractor also pulls links out of text, and for a quick flat list of everything it is the faster choice. This tool exists for the job that one does not do: grouping links by domain, counting them, stripping tracking parameters and catching bare addresses with no scheme. If you want the addresses, use the email tool; if you want to understand a page’s links, use this one.
For links stored inside a file rather than in text you can paste, the source has to be opened first. The PDF text extractor and the document text extractors produce plain text this tool then reads.
To understand the difference between extracting addresses and discovering them from scratch, read email extractor vs email finder.
URL format standards and extraction edge cases
A URL is a URI conforming to RFC 3986. This tool matches the http and
https schemes. The full path, query string and fragment are captured as written
in the source text.
Four edge cases: (a) a URL ending a sentence may carry a trailing full stop or comma —
the extractor strips trailing punctuation characters that cannot be a valid final URL
character (.,;!?)); (b) query strings containing an unencoded ampersand followed
by a word are captured in full, because extraction operates on raw characters rather than
parsed HTML entities; (c) IDN hostnames in Punycode (xn-- prefix) are returned
in Punycode form, not decoded to Unicode; (d) data: URIs are deliberately
excluded — their base64 payloads can be megabytes long and are not useful as navigable
links.
Frequently asked questions
How do I extract all the URLs from a block of text?
Paste the text into the box above and choose Every URL. Each link is pulled out, deduplicated and sorted, whether or not it was written with https:// in front.
Can it find links that do not start with http?
Yes. With Include links with no https:// on, it also catches www. addresses and bare domains like acme.com/blog, provided they end in a real top-level domain.
What does grouping by domain do?
Unique domains rolls every link up to one line per site, with a count of how many links point there. It turns a long list of URLs into a short list of the sites a page actually references.
What are tracking parameters and why remove them?
They are the utm_source, fbclid and gclid values that campaigns append to links to track clicks. They are not part of the real destination, so stripping them gives you the actual URLs rather than the campaign-tagged versions.
Does it upload my text?
No. The text is parsed inside your browser and nothing is transmitted, which is what makes it safe for a private email or an internal document.
Can it expand a bit.ly or other short link?
No. Expanding a shortened link means fetching it from a server, and this tool does not make network requests. The short link comes out as written.
What is the difference between this and the email extractor?
The email extractor gives a flat list of emails and links. This tool is built for links specifically: it groups them by domain, counts them, strips tracking parameters and catches addresses written without a scheme.