Why XML needs a path rather than a search
XML is a tree. The value you want is not at one location — it is at one location per element,
and the same tag name can appear at several depths with different meanings. A plain text search for
<title> finds the book titles, the chapter titles and the document title all mixed
together, with no way to tell them apart.
A path describes the shape once. book/title means "a title element whose parent is a book",
so chapter titles and the document title are left out. A bare title matches the tag wherever
it sits. The match is a suffix of each element's real path, which is why you only have to name the last
step or two rather than spell out the tree from the root.
How to extract values from XML
- Load your XML. Drop a .xml file or paste the text, then click Read the XML. Malformed XML is reported with the parser's own message and the point where it broke.
- Type a tag or a path. Type a single tag like title to match it anywhere, or a path like book/title to match only titles inside books. Prefix a name with @ for an attribute, e.g. @id or book/@id.
- Or leave it blank to flatten. With no field, every leaf element and attribute is listed as path and value — a whole tree turned into a two-column table you can scan or export.
- Export the values. Results are a path and value table. Copy the values, or download the full set as CSV or JSON with both columns.
What comes out
- Every matching element's text, one row per match, in document order — repeated tags list every occurrence, not just the first.
- The path of each value, so you can see exactly where in the tree it came from.
- Attribute values —
@idpulls every id attribute;book/@idpulls the id only from book elements. - A flattened view when the field is blank: every leaf and attribute as
path → value, which is the fastest way to understand an unfamiliar document. - Discovered tags, listed as clickable chips so you can pick a real path instead of guessing.
- CSV and JSON export carrying both the path and the value for every match.
What it accepts
Any well-formed XML: documents, RSS and Atom feeds, SVG, XHTML, sitemaps, configuration files and
API responses. CDATA sections are read as their text. Namespaced documents work either way — leave
Strip namespace prefixes on to match title whether the source wrote
title or dc:title, or turn it off to keep the prefix and match it exactly.
The on-page table shows up to 300 rows so a large feed still renders instantly; the CSV and JSON exports always contain every match. Pasted text is limited only by your device; dropped files are read up to 800,000 characters.
Parsing uses the browser's built-in XML parser, the same one your browser uses for every feed and document it loads. A file that opens in a browser will parse here.
Why the data never leaves your browser
Parsing and extraction run locally as JavaScript. Nothing you paste or drop is transmitted.
This matters for XML because the documents people need to inspect are often API responses, SOAP envelopes, configuration files and exports — the kind of payloads that carry API keys, connection strings, internal endpoints and customer records in the same tree as the one value you were after. Pasting one into an online XML viewer hands all of it to whoever runs that site.
What the path syntax does not do
The syntax is deliberately small — tag names separated by /, and @ for an
attribute. It is not XPath. It does not support:
- Predicates or filters. There is no
book[@in-stock='true']/title. Extract the values, then filter the result. - Positional selection.
book[1]is not supported; a path returns every match. - Wildcards or descendant axes. No
//or*. A bare tag already matches at any depth, which covers the common case. - Combining fields. One path at a time, so a two-column export means two passes — or leave the field blank and take the whole flattened tree at once.
For genuine XPath queries, a dedicated XML tool is the right choice. This covers the common "get me that value" case without one.
Who extracts values from XML
- Developers — pulling every URL out of a sitemap or every id out of an API response while debugging, without writing a throwaway parser.
- Feed and content work — collecting every link or title from an RSS or Atom feed.
- QA and support — checking which elements in a payload carry a value, and what it is.
- Data migration — flattening a system's XML export into a two-column list a spreadsheet accepts.
- Anyone handed an XML file who needs three values out of a thousand lines of tags.
XML compared with the other structured formats
XML nests, like JSON; the difference is only the syntax and that XML also carries attributes on
elements. The JSON field extractor is the sibling for JSON, using a
dot-and-[] path instead of a slash path. For a flat column of data, the
CSV column extractor selects a single column directly.
If your data is really an HTML table rather than data XML, the HTML table extractor turns it into rows and columns. XHTML and SVG, being XML, work here too.
Frequently asked questions
How do I get every value for one tag out of an XML file?
Load the file and type the tag name, for example title. Every element with that tag is listed with its path, in document order. Use book/title to match only titles inside books.
How do I extract an attribute rather than an element?
Prefix the name with @. Type @id to pull every id attribute, or book/@id to pull the id only from book elements. Keep Include attributes ticked.
What happens if I leave the field blank?
You get the whole document flattened: every leaf element and every attribute as a path and value pair. It is the quickest way to see what an unfamiliar XML file contains.
Does it handle namespaces like dc:title?
Yes. Leave Strip namespace prefixes on to match title whether the source wrote title or dc:title. Turn it off to keep the prefix and match dc:title exactly.
Is my XML uploaded anywhere?
No. Parsing happens in your browser — which matters, because XML payloads often carry API keys, connection strings and customer data alongside the value you wanted.
Why does my file say it is malformed?
XML must be well-formed: every tag closed, one root element, and & written as &. The tool shows the browser parser's own message and roughly where it stopped, which usually points straight at the problem.
Does it read CDATA sections?
Yes. The text inside a CDATA section is returned as the element's value, unescaped.
Is there a size limit?
Pasted text is limited only by your device. Dropped files are read up to 800,000 characters. The on-page table shows the first 300 matches; the CSV and JSON exports contain them all.