Extracts transactions from bank and credit card statement PDFs via OCR and outputs them as structured JSON — the first step towards reconciling supplier invoices with bank transactions.
Supported statements: Sparkasse checking account, Sparkasse Mastercard, American Express. OCR is done by a Google Document AI custom extractor; other providers (e.g. Mistral, Reducto) can be added behind the same interface later.
uv manages everything; there is no manual virtualenv handling.
- Put the Google service account key at
google-service-account.json(gitignored). - Copy
.env.exampleto.envand fill in the Document AI location and processor ID (.envis gitignored).
uv run extract.py sample_files/Example_Bank.PDF --model google --source sparkasse-bank
uv run extract.py sample_files/Example_MasterCard.PDF --model google --source sparkasse-mastercard
uv run extract.py sample_files/Example_AMEX.pdf --model google --source amex --statement-date 2026-06-14
uv run extract.py sample_files/Example_Bank.PDF --model reducto --source sparkasse-bank--modelpicks the OCR provider (googleorreducto).--sourcepicks the statement-specific normalization rules.--statement-datesupplies the year for statements that print dates without one (AMEX).--output transactions.jsonwrites to a file instead of stdout.--save-raw raw.jsonstores the raw OCR result;--from-raw raw.jsonre-normalizes from it without paying for another OCR call. The Document AI extractor is an ML model and not perfectly deterministic across runs — for reproducible reruns, work from a saved raw file.--model reductouses Reducto Extract. It readsREDUCTO_API_KEYfrom the environment/.env, and can fall back to the MCP login key written byuvx mcp-server-reducto --loginat~/.reducto/config.yaml.
uv run invoice-ocr <args> is an equivalent installed entry point.
The output JSON contains transactions (typed, classified, with stable
deterministic IDs) and errors (rows that could not be normalized).
To check which debit transactions already have a supplier invoice, put the invoice PDFs in one folder and run:
uv run invoice-match \
--transactions bank_2026_06_reducto.json \
--invoice-dir invoices \
--output invoice_match_report.json \
--save-invoices extracted_invoices.json \
--parallelism 4invoice-match extracts every PDF in --invoice-dir with Reducto, matches
the extracted invoice totals against debit transactions by exact EUR amount
and merchant/vendor similarity, and writes a report with:
matched_transactionstransactions_without_invoiceinvoices_without_transactioninvoice_extraction_errors
Invoice extraction uses up to 4 parallel Reducto requests by default. Tune this
with --parallelism 2 or --parallelism 6 depending on API limits. Extracted
invoices are also cached by PDF hash in .invoice_cache/invoices.json, so
rerunning with the same PDFs skips Reducto automatically. Use --no-invoice-cache
to force a fresh extraction or --invoice-cache path/to/cache.json to choose a
different cache file.
When iterating on matching rules, reuse the saved invoice summaries so no documents are uploaded again:
uv run invoice-match \
--transactions bank_2026_06_reducto.json \
--invoice-dir invoices \
--from-invoices extracted_invoices.json \
--output invoice_match_report.jsonFor a local test UI:
uv run invoice-uiThen open http://127.0.0.1:8765. The UI supports statement uploads, per-PDF
statement source selection, invoice uploads that can be added and removed over
multiple picker actions, model selection, cached invoice JSONs, and a
parallelism setting for invoice extraction.
It also provides:
- terminal debug logging per run via a UI checkbox
- JSON/CSV downloads for reports, matches, missing transactions, and invoice cache data
- manual matching from an unmatched transaction to an unmatched invoice
- confidence badges and duplicate/low-confidence warnings
- local run folders under
runs/when "Lauf lokal speichern" is enabled
PDF ──(providers.py: Document AI custom extractor)──▶ RawTransaction*
──(normalize.py: per-source sign/date/classification rules)──▶ BankTransaction*
invoice_ocr/providers.py— OCR providers. Each returns provider-agnosticRawTransactionrows (date, description, amounts as raw text).invoice_ocr/normalize.py— turns raw rows intoBankTransactions: parses German/English amount and date formats, applies each source's sign convention, and classifies rows (supplier payment, card purchase, FX fee, settlement, balance line, refund).invoice_ocr/models.py— Pydantic models; money isDecimal, never float.invoice_ocr/cli.py— the command line interface.
The google provider uses a custom extraction processor
(amalytix-statement-extractor), a Document AI processor trained on these
statement layouts to recognize whole transactions. Unlike plain OCR — which
returns text lines and half-detected tables that would have to be
reassembled with fragile heuristics — it directly emits one transaction
entity per statement row with these properties:
| Property | Example |
|---|---|
date |
03.06.26 |
description |
OPENAI *CHATGPT SUBSCR |
amount_eur |
17,22 - |
original_amount |
20,00 (optional) |
original_currency |
USD (optional) |
One extraction is one synchronous process_document call: the PDF bytes are
sent to https://{location}-documentai.googleapis.com for the processor
projects/{project}/locations/{location}/processors/{processor_id}, using
the service account key for authentication (the project ID is read from that
file unless GOOGLE_PROJECT_ID overrides it). Each call is billed per page,
so cache results with --save-raw when iterating.
providers.py then maps every transaction entity to a RawTransaction,
keeping all values as raw text — interpreting signs, date formats, and
currencies is deliberately left to normalize.py, where the per-bank rules
live. Entities without a date or amount are skipped with a warning, as are
entity types other than transaction. Because the processor is an ML model,
extraction is not perfectly deterministic and property values can contain
OCR noise; anything that fails to normalize later lands in the output's
errors list rather than being silently dropped.
The Google Cloud project also contains pretrained amalytix-bank
(BANK_STATEMENT_PROCESSOR), amalytix-invoice (INVOICE_PROCESSOR), and
amalytix-ocr (OCR_PROCESSOR) processors — switch via GOOGLE_PROCESSOR_ID
in .env, though only the custom extractor emits the transaction entities
this pipeline consumes.
uv run pytest # fast tests (no network)
uv run pytest -m requires_google # live OCR smoke tests, costs money
uv run ruff check .
uv run mypySample statements live in sample_files/ (gitignored — they contain real
account data).