First, what kind of PDF do you have?
"PDF" covers three very different things, and the right approach depends on which one is on your desk. Open a typical file and try to select a line of text.
- KindDigital (born digital)
- How to tellText selects and copies cleanly.
- What it meansThe words are already inside the file. Extraction is about finding the right values, not reading them.
- KindScanned (image only)
- How to tellNothing selects, or the whole page selects as one picture.
- What it meansIt needs OCR first. Scan quality now matters as much as the extraction method.
- KindFillable form
- How to tellBoxes you can click and type into.
- What it meansThe values may be stored as form fields you can export directly, with no reading at all.
| Kind | How to tell | What it means |
|---|---|---|
| Digital (born digital) | Text selects and copies cleanly. | The words are already inside the file. Extraction is about finding the right values, not reading them. |
| Scanned (image only) | Nothing selects, or the whole page selects as one picture. | It needs OCR first. Scan quality now matters as much as the extraction method. |
| Fillable form | Boxes you can click and type into. | The values may be stored as form fields you can export directly, with no reading at all. |
Most real piles are mixed: digital invoices emailed by large vendors, scans from small ones, and the occasional phone photo saved as a PDF. Plan for the worst of them, not the best.
Decide what "structured" means for you
Before comparing tools, write down five things. They decide the answer more than any feature list.
- The fields. The exact values you need, for example vendor, invoice number, date, PO number and total. Fewer is better.
- Tables. Whether you need line items (description, quantity, unit price, amount) or only header values. Line items are much harder.
- Variety. How many different layouts you receive. Five senders is a different problem from five hundred.
- Volume and timing. Twenty a month, or two thousand a day, and whether data is needed in minutes or by month end.
- The destination. A spreadsheet, an accounting import, a document library, or another system through an API.
Four approaches, compared honestly
- ApproachManual entry
- Good atNo setup, any document, a person understands context.
- Breaks whenVolume grows. Errors rise with fatigue, and it does not scale.
- FitsA few dozen documents a month.
- ApproachCopy, paste and PDF-to-text tools
- Good atFree or cheap for digital PDFs.
- Breaks whenScans, tables (columns scramble) and anything repeated daily.
- FitsOne-off jobs on digital files.
- ApproachTemplates and zonal OCR
- Good atFixed forms that never change: your own forms, government forms.
- Breaks whenA sender moves a field, adds a line or sends a new layout. Every layout is a template to keep up.
- FitsFew layouts, high volume, stable forms.
- ApproachLearned extraction
- Good atMany layouts. Learns fields and tables from examples and corrections.
- Breaks whenYou skip review entirely. Good tools tell you when they are unsure; use that.
- FitsMixed documents from many senders.
| Approach | Good at | Breaks when | Fits |
|---|---|---|---|
| Manual entry | No setup, any document, a person understands context. | Volume grows. Errors rise with fatigue, and it does not scale. | A few dozen documents a month. |
| Copy, paste and PDF-to-text tools | Free or cheap for digital PDFs. | Scans, tables (columns scramble) and anything repeated daily. | One-off jobs on digital files. |
| Templates and zonal OCR | Fixed forms that never change: your own forms, government forms. | A sender moves a field, adds a line or sends a new layout. Every layout is a template to keep up. | Few layouts, high volume, stable forms. |
| Learned extraction | Many layouts. Learns fields and tables from examples and corrections. | You skip review entirely. Good tools tell you when they are unsure; use that. | Mixed documents from many senders. |
Manual entry is not always wrong
If you process thirty documents a month and they go into one system, typing is honest and cheap. The case for software starts when the volume, the variety or the cost of mistakes grows, or when the same people keep retyping the same kinds of values.
Templates are reliable until they are not
Zonal OCR reads a value from a fixed box on the page. On a form you control, that works well. On invoices or freight bills from hundreds of senders it turns into a maintenance job: every new layout needs a template, and every layout change quietly breaks one.
Learned extraction needs a feedback loop
Learned tools work out where the fields are from a set of examples, then improve from the corrections people make. The two things to insist on: they must say how sure they are and why a document needs a look, and corrections must actually make the next documents better.
Desktop software or a cloud service?
Separate from how the data is extracted is where it happens. Neither is better in general; they suit different teams.
- QuestionWhere documents are read
- Local desktop softwareOn a PC you control
- Cloud service or APIOn the provider’s servers
- QuestionPaper scanning
- Local desktop softwareNatural fit: scanner on the same PC
- Cloud service or APIPossible, but documents are uploaded first
- QuestionDocuments arriving by email or from systems
- Local desktop softwareSaved to a folder, then processed
- Cloud service or APINatural fit: send them straight in
- QuestionGetting data into other systems
- Local desktop softwareFiles and folders, or connected destinations
- Cloud service or APIDownload, SFTP or webhook into your workflow
- QuestionIT and privacy reviews
- Local desktop softwareSimpler: images are not sent out to be read
- Cloud service or APIReview the provider and its hosting
- QuestionCapacity
- Local desktop softwareLimited by the PC
- Cloud service or APIGrows with the service
| Question | Local desktop software | Cloud service or API |
|---|---|---|
| Where documents are read | On a PC you control | On the provider’s servers |
| Paper scanning | Natural fit: scanner on the same PC | Possible, but documents are uploaded first |
| Documents arriving by email or from systems | Saved to a folder, then processed | Natural fit: send them straight in |
| Getting data into other systems | Files and folders, or connected destinations | Download, SFTP or webhook into your workflow |
| IT and privacy reviews | Simpler: images are not sent out to be read | Review the provider and its hosting |
| Capacity | Limited by the PC | Grows with the service |
A worked example: one invoice, start to finish
Here is a fictional sample invoice as CapturePoint 6 extracts it. The fields read from the header were Northbridge Office Supply, invoice INV-2401, dated 09/04/2026, due 10/04/2026, PO-1031, total 480.00. The table has four lines.
Look at line 1. Its amount was captured as 141.00, but 4 times 35.00 is 140.00. Rather than pass that on, the line is flagged: "Line 1 does not add up. Correct it, or mark it right as printed." This is the kind of check that separates extraction you can trust from extraction you have to re-check by hand.
Once confirmed, the invoice leaves as a searchable PDF (optionally PDF/A), a data file with every captured value and a text file, named and filed by the fields you choose.
How to evaluate any extraction tool
Demos use clean documents. Your decision should be based on yours. Run every candidate through the same test.
An evaluation you can run in a week
- Collect 50 to 100 real documents, including the worst: faxes, skewed scans, multi-page invoices, new senders.
- Write down the correct value of every field you need for each one before you start. That is your answer key.
- Measure field by field, not document by document. A document with one wrong total is a wrong document.
- Count how often the tool says it is unsure, and how often it is wrong without saying so. The second number matters most.
- Test line items separately: row counts, columns and whether rows add up to the total.
- Include a stack scanned as one file and check that it is split into the right documents.
- Time a person reviewing 25 documents. Review speed is where the real cost is.
- Export to your real destination and confirm the next system takes it without retyping.
- Correct a few mistakes and run similar documents again. Did it learn?
When Paige fits, and when CapturePoint 6 does
Ademero makes one of each kind, so we can be straight about which suits you. Both learn from a few examples, split stacks into documents, sort them by type, and read fields and line items.
- What it is
- PaigeA cloud service: send in documents, get clean data back
- CapturePoint 6A Windows 10/11 (64-bit) app on your own PC
- Where reading happens
- PaigeOn Google Cloud
- CapturePoint 6Locally on the PC, graphics card optional
- How documents get in
- PaigeSent to the service
- CapturePoint 6TWAIN scanners, or PDF, TIFF, JPEG, PNG, BMP and GIF files
- How results come out
- PaigeDownload, SFTP or webhook
- CapturePoint 6Folders (searchable PDF, data file, text file) or Content Central; also SharePoint/OneDrive, Google Drive, Dropbox or Nucleus One
- Best for
- PaigeDocuments that arrive digitally and should flow into other systems
- CapturePoint 6Paper scanned in-house, and teams that want documents read on their own PCs
- How to try it
- PaigeStart free on paige.app
- CapturePoint 6Free trial starts on first launch, no form
| Paige | CapturePoint 6 | |
|---|---|---|
| What it is | A cloud service: send in documents, get clean data back | A Windows 10/11 (64-bit) app on your own PC |
| Where reading happens | On Google Cloud | Locally on the PC, graphics card optional |
| How documents get in | Sent to the service | TWAIN scanners, or PDF, TIFF, JPEG, PNG, BMP and GIF files |
| How results come out | Download, SFTP or webhook | Folders (searchable PDF, data file, text file) or Content Central; also SharePoint/OneDrive, Google Drive, Dropbox or Nucleus One |
| Best for | Documents that arrive digitally and should flow into other systems | Paper scanned in-house, and teams that want documents read on their own PCs |
| How to try it | Start free on paige.app | Free trial starts on first launch, no form |
Neither is the answer for a one-off batch of twenty digital PDFs (copy them by hand) or for a fillable form whose values you can export directly. Help for each is in the Ademero help library: help.ademero.com/paige and help.ademero.com/capturepoint.
Specific documents
- Freight paperwork: capturing data from bills of lading.
- Vendor tax forms: getting vendor data off W-9 forms.
- Claims mail: indexing claims by claim and policy number.