Power Query can't read my PDFs: alternatives that work (2026)
Power Query's PDF connector imports the tables it detects in a PDF, so it fails in three common cases: the "From PDF" option is missing from your Excel, the PDFs do not share the same layout, or the PDFs are scans with no text. The first is a version issue, the third needs OCR, and the second needs a method that finds values by label or pattern instead of by table position, such as a script, a template-based parser, or a browser tool that extracts named fields from each PDF.
Which failure do you have?
| Symptom | Likely cause | What to do |
|---|---|---|
| No "From PDF" under Get Data > From File | Excel for Mac, or a one-time-purchase or older Windows version | Use Excel for Microsoft 365 on Windows, or one of the alternatives below |
| Navigator shows no tables, or only "Page001" | The PDF has no text layer (scan) or no table structure | OCR the scan first; for non-table values use label-based extraction |
| Combining a folder gives shifted or mixed columns | Files have different layouts; the steps were built from one example file | Extract by label per file, or split the folder by layout |
| Values land in one merged column | Table detection merged cells | Try the EnforceBorderLines or Implementation options of Pdf.Tables, or extract by label |
| Refresh is slow or times out on long PDFs | Large documents processed in one pass | Limit pages with StartPage and EndPage |
Why different layouts break the query
When you combine files from a folder, Power Query analyses an example file, "by default, the first file in the list", builds the extraction steps on it, then applies those steps to every file. Microsoft states that combining works "as long as they have the same file type and structure (including the same columns)" (Combine files overview). Invoices from ten suppliers are ten structures. The connector itself returns "any tables found" (PDF connector), and header fields such as the invoice number or the issue date are often not in a table at all.
The practical consequence: either the source PDFs share one structure, or each layout needs its own query or custom M code.
Fixes inside Power Query
- One query per layout: group PDFs by supplier into subfolders, build one combined query per folder, then append the results.
- Pdf.Tables options: the connector accepts
StartPage,EndPage,MultiPageTables,EnforceBorderLinesandImplementation(documentation). Changing them can fix merged or split tables. - Select columns by name: use
Table.SelectColumnswithMissingField.UseNullso a file with a missing column still fits the schema.
These keep you in Excel but require M code and a query per layout.
Alternatives that do not depend on layout
| Option | How it finds values | Cost | Data location |
|---|---|---|---|
| Python script (pdfplumber or pypdf + regex) | Your patterns | Your time | Local |
| Docparser | Parsing rules per layout | From $39/month, monthly billing (pricing) | Vendor servers |
| Parseur | Templates or AI | Free up to 20 pages/month, then from 49 €/month (pricing) | Vendor servers |
| Nanonets | AI extraction workflows | Pay per processing block, $50 free credits (pricing) | Vendor servers |
| Browser extractor by label, regex or detector | Labels, regex, detectors for dates, totals, VAT IDs, IBAN | Free up to 5 PDFs per export, then $4.99 for 24 h | Your browser |
Prices as shown on the vendors' pages in October 2026.
The browser option is the PDF data extractor: drop the folder of PDFs, keep the suggested fields or add your own ("Total Due", "Invoice Date"), check the full table on screen (empty cells in orange, editable), then download XLSX with real dates and numbers, or CSV. Nothing is uploaded, and field setups are saved as reusable templates.
FAQ
Does Excel for Mac have the PDF connector? Microsoft Q&A answers from 2023 to 2025 report that it does not (thread). Check your version under Data > Get Data.
Can Power Query read scanned PDFs? No. It needs a text layer. Run OCR first, then import.
Why does my combined query work for the first file only? The steps were recorded on the example file. Files with a different layout produce different tables, so later steps reference columns that do not exist there.
Is there a way to get just the invoice number and total from each PDF? Yes: extract fields by label or pattern, one row per file, instead of importing tables. A script or a label-based extractor does this regardless of where the value sits on the page.