How to extract the same fields from many PDFs into Excel (2026)

Updated · 5 min read

To pull the same fields out of a folder of PDFs into one Excel sheet, you need a method that finds each value by its label or format (for example the amount after "Total Due", or the first date after "Invoice Date") rather than by its position on the page. Excel's Power Query works when every PDF has the same table layout; when layouts differ, use a template-based parser, a script, or a browser tool that searches each PDF's text by label and gives you one row per file. Scanned PDFs need OCR first, whatever the method.

What "the same fields" means in practice

The typical job: 40 supplier invoices, 12 monthly statements or 200 reports, and you want one spreadsheet with a row per document and columns such as file name, document number, date, total, VAT ID, customer reference. The values exist in every PDF, but not in the same place: one supplier puts the total at the bottom right, another in a summary box at the top, a third writes "Amount due" instead of "Total".

That is why copy-paste works for five files and breaks down at fifty, and why tools built to import tables struggle: the fields you want are usually not in a table.

The methods compared

MethodHandles different layoutsDates and amounts as real Excel valuesSetup effortWhere the PDFs go
Manual copy-pasteYesOnly if you retype carefullyNone, but scales badlyStays on your computer
Excel Power Query, PDF connector + FolderPoorly: it imports detected tables, built from one sample fileYes, after type conversion stepsMedium (queries, M code)Stays on your computer
Acrobat "Export PDF" to Excel, then formulasOne file at a time, layout kept as-isPartlyHigh for many filesDepends on your Acrobat setup
Script (Python with pdfplumber or pypdf + regex)Yes, if you write the rulesYesHigh (code and maintenance)Stays on your computer
Hosted parsing services (Docparser, Parseur, Nanonets)Yes, with templates or AI per layoutYesMediumUploaded to the vendor
Browser tool that extracts by label, regex or detectorYesYesLowStays in your browser

Method 1: Power Query (Get Data > From PDF, or From Folder)

Microsoft's PDF connector "returns any tables found in a PDF file". To process several PDFs, Microsoft points to the Folder connector and the combine files feature, which builds its steps from an example file (by default, the first file in the list) and works "as long as they have the same file type and structure (including the same columns)".

It suits a folder of statements from a single bank, all generated by the same system. It struggles with invoices from many suppliers: the tables detected differ from file to file, and header fields such as the invoice number are often outside any table. Microsoft Q&A threads also report that the PDF connector is not available in Excel for Mac.

Method 2: a script

A Python script reads each PDF's text, applies one regular expression per field, and writes a CSV. It handles any layout you write a rule for and keeps files local. The cost is writing and maintaining the patterns, handling number formats (1,234.56 vs 1.234,56), and dates (03/04/2026 is March 4 in the US and April 3 in Europe).

Method 3: hosted parsers

Docparser, Parseur and Nanonets let you define fields once per layout (or use AI extraction) and export to spreadsheets or integrations. They are subscriptions aimed at recurring volumes. On their pricing pages in October 2026, Docparser starts at $39 per month for 100 credits on monthly billing, Parseur has a free tier of 20 pages per month and a first paid tier at 49 € per month for 100 pages, and Nanonets charges per processing block with $50 of free credits. Your PDFs are uploaded to their servers.

Method 4: extract by label in the browser

The PDF data extractor reads the text of each PDF in your browser and builds one row per file. Each column is a field defined by a label ("Invoice No", value on the same line or the line below), a regular expression, or a ready-made detector (date, total, subtotal, tax, EU VAT ID, IBAN with checksum, email, phone, invoice number). It suggests label: value pairs from your first PDF, flags empty cells in orange and lets you correct them, and exports XLSX with real dates and numbers, or CSV. The preview is complete; export is free up to 5 PDFs, then $4.99 for 24 hours. Field setups are saved as templates for next month.

Checks before you trust the spreadsheet

  • Row count: one row per PDF. A missing row usually means an unreadable or password-protected file.
  • Empty cells: each one is a field the method could not find. Open that PDF's text to see why (different label, value on another line).
  • Totals: sum the total column and compare it with the sum you expect (bank deposits, ledger balance).
  • Duplicates: the same invoice number twice usually means the same PDF was saved twice.
  • Scanned files: a PDF with no selectable text produces nothing until it goes through OCR.

FAQ

Can Excel do this without add-ins? Yes when the PDFs share a layout: Get Data > From File > From Folder, then combine the PDFs. When layouts differ, the combined query mixes or drops columns and needs manual M code.

What about scanned invoices? Every text-based method needs a text layer. Run the scans through OCR (Acrobat "Recognize text", or a scanner set to searchable PDF), then extract.

How do I keep dates and amounts as numbers? Normalise them before they reach Excel: dates to YYYY-MM-DD, amounts with a single decimal separator. In XLSX, store them as date and number cells, not text.

Is it safe to upload invoices to an online converter? Invoices carry bank details and customer data. Hosted services process them on their servers; check their retention and data processing terms. Local scripts and in-browser tools do not send the files anywhere.

Extract data from many PDFs

  • PDF
  • Excel
  • CSV

Export free up to 5 PDFs, then $4.99

Extract data from PDFs