All PDFind guides

Work with scanned PDFs in Excel

PDFind uses English and Polish OCR to recognize printed text in scanned PDFs. You can then select values or tables and insert them into your workbook.

Digital PDF or scan?

A digital PDF usually contains embedded text that PDFind can read directly. A scan contains a picture of the page, so its characters need to be recognized first. Recognition quality depends on the original image.

Extract and check the data

  1. Open the PDF in PDFind and go to the page you need.
  2. Allow the first recognition operation to finish; it can take longer than reading embedded text.
  3. Select the destination cell in Excel, then select the value or use Select a table.
  4. Compare the inserted result with the PDF image before using it in calculations or reports.

What to check carefully

Use a clear source

Low-resolution, rotated, faint or noisy scans need more review. Clean printed text works better than handwriting. PDFind does not guarantee handwriting recognition or reconstruction of merged cells and complex multi-page table headers.

Where OCR runs

Recognition runs inside the Office task pane. PDF content is not sent to an external OCR provider or to PDFind's licensing and telemetry API. Saved source documents use Microsoft storage; Microsoft sign-in is required. Read the security details.

For other file restrictions, see supported PDFs and known limits. To verify an imported value later, follow the source-review guide.