Skip to content

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Latest commit

 

History

11 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

invoice-ocr

Extracts transactions from bank and credit card statement PDFs via OCR and outputs them as structured JSON — the first step towards reconciling supplier invoices with bank transactions.

Supported statements: Sparkasse checking account, Sparkasse Mastercard, American Express. OCR is done by a Google Document AI custom extractor; other providers (e.g. Mistral, Reducto) can be added behind the same interface later.

Setup

uv manages everything; there is no manual virtualenv handling.

  1. Put the Google service account key at google-service-account.json (gitignored).
  2. Copy .env.example to .env and fill in the Document AI location and processor ID (.env is gitignored).

Usage

uv run extract.py sample_files/Example_Bank.PDF --model google --source sparkasse-bank
uv run extract.py sample_files/Example_MasterCard.PDF --model google --source sparkasse-mastercard
uv run extract.py sample_files/Example_AMEX.pdf --model google --source amex --statement-date 2026-06-14
uv run extract.py sample_files/Example_Bank.PDF --model reducto --source sparkasse-bank
  • --model picks the OCR provider (google or reducto).
  • --source picks the statement-specific normalization rules.
  • --statement-date supplies the year for statements that print dates without one (AMEX).
  • --output transactions.json writes to a file instead of stdout.
  • --save-raw raw.json stores the raw OCR result; --from-raw raw.json re-normalizes from it without paying for another OCR call. The Document AI extractor is an ML model and not perfectly deterministic across runs — for reproducible reruns, work from a saved raw file.
  • --model reducto uses Reducto Extract. It reads REDUCTO_API_KEY from the environment/.env, and can fall back to the MCP login key written by uvx mcp-server-reducto --login at ~/.reducto/config.yaml.

uv run invoice-ocr <args> is an equivalent installed entry point.

The output JSON contains transactions (typed, classified, with stable deterministic IDs) and errors (rows that could not be normalized).

Invoice matching

To check which debit transactions already have a supplier invoice, put the invoice PDFs in one folder and run:

uv run invoice-match \
  --transactions bank_2026_06_reducto.json \
  --invoice-dir invoices \
  --output invoice_match_report.json \
  --save-invoices extracted_invoices.json \
  --parallelism 4

invoice-match extracts every PDF in --invoice-dir with Reducto, matches the extracted invoice totals against debit transactions by exact EUR amount and merchant/vendor similarity, and writes a report with:

  • matched_transactions
  • transactions_without_invoice
  • invoices_without_transaction
  • invoice_extraction_errors

Invoice extraction uses up to 4 parallel Reducto requests by default. Tune this with --parallelism 2 or --parallelism 6 depending on API limits. Extracted invoices are also cached by PDF hash in .invoice_cache/invoices.json, so rerunning with the same PDFs skips Reducto automatically. Use --no-invoice-cache to force a fresh extraction or --invoice-cache path/to/cache.json to choose a different cache file.

When iterating on matching rules, reuse the saved invoice summaries so no documents are uploaded again:

uv run invoice-match \
  --transactions bank_2026_06_reducto.json \
  --invoice-dir invoices \
  --from-invoices extracted_invoices.json \
  --output invoice_match_report.json

For a local test UI:

uv run invoice-ui

Then open http://127.0.0.1:8765. The UI supports statement uploads, per-PDF statement source selection, invoice uploads that can be added and removed over multiple picker actions, model selection, cached invoice JSONs, and a parallelism setting for invoice extraction.

It also provides:

  • terminal debug logging per run via a UI checkbox
  • JSON/CSV downloads for reports, matches, missing transactions, and invoice cache data
  • manual matching from an unmatched transaction to an unmatched invoice
  • confidence badges and duplicate/low-confidence warnings
  • local run folders under runs/ when "Lauf lokal speichern" is enabled

How it works

PDF ──(providers.py: Document AI custom extractor)──▶ RawTransaction*
    ──(normalize.py: per-source sign/date/classification rules)──▶ BankTransaction*
  • invoice_ocr/providers.py — OCR providers. Each returns provider-agnostic RawTransaction rows (date, description, amounts as raw text).
  • invoice_ocr/normalize.py — turns raw rows into BankTransactions: parses German/English amount and date formats, applies each source's sign convention, and classifies rows (supplier payment, card purchase, FX fee, settlement, balance line, refund).
  • invoice_ocr/models.py — Pydantic models; money is Decimal, never float.
  • invoice_ocr/cli.py — the command line interface.

Google Document AI extraction

The google provider uses a custom extraction processor (amalytix-statement-extractor), a Document AI processor trained on these statement layouts to recognize whole transactions. Unlike plain OCR — which returns text lines and half-detected tables that would have to be reassembled with fragile heuristics — it directly emits one transaction entity per statement row with these properties:

Property Example
date 03.06.26
description OPENAI *CHATGPT SUBSCR
amount_eur 17,22 -
original_amount 20,00 (optional)
original_currency USD (optional)

One extraction is one synchronous process_document call: the PDF bytes are sent to https://{location}-documentai.googleapis.com for the processor projects/{project}/locations/{location}/processors/{processor_id}, using the service account key for authentication (the project ID is read from that file unless GOOGLE_PROJECT_ID overrides it). Each call is billed per page, so cache results with --save-raw when iterating.

providers.py then maps every transaction entity to a RawTransaction, keeping all values as raw text — interpreting signs, date formats, and currencies is deliberately left to normalize.py, where the per-bank rules live. Entities without a date or amount are skipped with a warning, as are entity types other than transaction. Because the processor is an ML model, extraction is not perfectly deterministic and property values can contain OCR noise; anything that fails to normalize later lands in the output's errors list rather than being silently dropped.

The Google Cloud project also contains pretrained amalytix-bank (BANK_STATEMENT_PROCESSOR), amalytix-invoice (INVOICE_PROCESSOR), and amalytix-ocr (OCR_PROCESSOR) processors — switch via GOOGLE_PROCESSOR_ID in .env, though only the custom extractor emits the transaction entities this pipeline consumes.

Development

uv run pytest                      # fast tests (no network)
uv run pytest -m requires_google   # live OCR smoke tests, costs money
uv run ruff check .
uv run mypy

Sample statements live in sample_files/ (gitignored — they contain real account data).

About

No description, website, or topics provided.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages