Why merged cells break naive conversion
An HTML table is not a grid. It is a list of rows, each holding a list of cells, and any cell may claim
several columns with colspan or several rows with rowspan. A row that looks
eight columns wide on screen may contain only five cells in the markup.
Converting row by row therefore produces a CSV where every row after a merged cell is shifted left, which is exactly the failure people see when they paste a web table into Excel. This tool builds a dense grid first: a cell spanning three columns is written into three positions, and a cell spanning two rows is written into the row below as well. The result has the same shape as the rendered table.
How to convert an HTML table to CSV
- Copy the HTML. On the page, either right-click and choose View page source and copy the whole thing, or open developer tools, find the <table> element, right-click it and copy the outer HTML. Either works — the whole page is fine.
- Paste it and find the tables. Every table in the markup is listed with its size, and its caption when it has one, so picking the right one on a page full of tables is straightforward.
- Check the preview. Merged cells are already expanded at this point, so what you see is what the CSV will contain.
- Download. Export the selected table as CSV, all tables in one file, or copy as tab-separated values to paste straight into a spreadsheet.
What comes out
- Every table in the markup, listed with its dimensions and caption.
- A rectangular grid with
colspanandrowspanexpanded, so columns line up. - Text only — links, bold, spans and nested markup are reduced to their text, with whitespace collapsed.
- Properly quoted CSV, so cells containing commas or quotes survive the trip into a spreadsheet.
- All tables in one file, each under a labelled heading, for pages carrying several.
What it accepts
Paste a full page source, a fragment, or a single copied <table> element. The
markup is parsed with the browser's own HTML parser, so unclosed tags, missing
<tbody> elements and the other everyday malformations of real-world HTML are resolved
exactly as a browser resolves them — which is the behaviour you want, since the browser's interpretation
is the table you actually saw.
Nested tables are read as separate tables, which is usually right: old layouts used tables for page structure, so the outer one is scaffolding and the inner one is your data.
Why the page content stays with you
Parsing happens in your browser. Nothing you paste is transmitted.
The tables worth extracting are frequently behind a login — an admin dashboard, a bank statement, a CRM report, an internal wiki. Pasting that markup into a server-side converter sends the data, and often session identifiers embedded in the surrounding HTML, to a third party. Local parsing avoids both.
What it cannot read
The most common disappointment: many modern sites have no <table> at
all. Data laid out with divs and CSS grid looks exactly like a table on screen but is not one in
the markup, and this tool will correctly report finding nothing. For those, select the rendered table in
the page and paste it into a spreadsheet, which usually preserves the columns.
Three further limits:
- Merged cells are duplicated, not merged. A cell spanning three columns appears three times. That keeps the grid rectangular, which CSV requires.
- Tables built by JavaScript need the rendered HTML. View page source shows the original markup; copy the element from developer tools instead.
- Formatting is discarded. Colours, links and images are not carried into CSV — only the text.
Who converts HTML tables
- Analysts — lifting a published data table off a page and into a spreadsheet.
- Researchers — collecting reference tables from documentation or statistical sites.
- Developers — turning a rendered report back into data for testing.
- Anyone copying from an admin panel where there is no export button.
- Migration work — extracting tables from an old site before it is retired.
HTML tables compared with PDF tables
An HTML table declares its own structure — rows and cells are explicit — so conversion is exact apart from merged cells. A table in a PDF has no structure at all, only text at coordinates, so the PDF table extractor has to infer columns from the gaps between words and sometimes gets it wrong. Given the choice, always take the HTML.
If the data began as a spreadsheet, the Excel data extractor is better still. Once you have CSV, the CSV column extractor narrows it to the columns you need.
HTML table format and extraction edge cases
An HTML table is defined by the HTML Living Standard: <table>,
<thead>, <tbody>, <tfoot>,
<tr>, <th> and <td>. The
colspan and rowspan attributes extend cells across columns and
rows; the extractor expands them by repeating the value in every covered cell, which matches
the browser rendering but widens the column count when heavy spanning is used.
Three edge cases: (a) a nested table — a <table> inside a
<td> — is extracted as a separate table, not as a nested value inside the
parent cell; (b) tables generated by JavaScript are only present if you paste the rendered
HTML (use the browser's developer tools to copy the element's outer HTML), not the original
page source of an SPA; (c) CSS-styled <div> grids that look like tables
visually are not <table> elements and are not returned — only genuine
HTML table markup is parsed.
Frequently asked questions
How do I convert an HTML table to CSV?
Paste the page's HTML, pick the table from the list and click Download this table as CSV. Merged cells are expanded so the columns line up.
Where do I get the HTML from?
Right-click the page and choose View page source, or open developer tools, right-click the <table> element and copy its outer HTML. Pasting the whole page is fine.
It says no table was found, but I can see one.
The data is probably laid out with divs and CSS rather than a real table element, which is common on modern sites. Select the table on the page and paste it into a spreadsheet instead.
Is the page content uploaded?
No. Parsing happens in your browser, which matters because the tables worth extracting are often behind a login.
How are merged cells handled?
A cell spanning three columns is written into all three positions, and a rowspan is repeated into the rows below. That keeps the grid rectangular, which is what CSV requires.
Can it handle a page with several tables?
Yes. Every table is listed with its size and caption, and you can export one or all of them.
What about tables generated by JavaScript?
View page source shows the original markup, which will not contain them. Copy the element from developer tools, which shows the rendered result.
Are links and images kept?
No — CSV holds text. Link text is kept; the URL and any images are not.