Skip to content
Merged
104 changes: 104 additions & 0 deletions examples/fan_out_provenance.py
Original file line number Diff line number Diff line change
@@ -0,0 +1,104 @@
"""Fan-out ingestion, and the origin DataJoint records for it.

One ``make()`` parses a single source file into several entry-point tables that
carry no foreign key back to it. The dependency graph cannot record those
writes -- that is what makes the pattern a deliberate exception -- so the
question "where did this Subject row come from" has nothing structural to
answer it.

Since 2.3.4 the framework answers it anyway. Every row written into a Manual
table carries a hidden ``_prov`` attribute, and a row written from inside an
ingesting ``make()`` records the ingesting table and key: exactly the link the
absent foreign key would have carried. Nothing below asks for that, and nothing
can forge it.

Run with a configured source, the way a deployment would set it::

DJ_PROVENANCE_SOURCE='{"system": "AcquisitionShare", "root": "/mnt/raw"}' \
python examples/fan_out_provenance.py
"""

import datajoint as dj

schema = dj.Schema("fan_out_provenance_demo")


@schema
class RecordingFile(dj.Manual):
"""The source record: one row per file the pipeline has been told about."""

definition = """
file_id : int32
---
path : varchar(255)
"""


@schema
class Subject(dj.Manual):
definition = """
subject_id : int32
---
species : varchar(64)
source_file : int32 # the RecordingFile this row was parsed from
"""


@schema
class Session(dj.Manual):
definition = """
session_id : int32
---
session_date : date
source_file : int32 # the RecordingFile this row was parsed from
"""


@schema
class Ingest(dj.Imported):
"""Parses one file into several entry-point tables that do not depend on it."""

definition = """
-> RecordingFile
---
n_entities : int32
"""

def make(self, key):
meta = parse(key["file_id"])

# The fan-out. Subject and Session have no foreign key back to Ingest,
# so dj.Diagram renders them as unconnected nodes. `source_file` is the
# link the pipeline itself queries; `_prov` is the audit record, written
# by the framework with this table and key in it.
Subject.insert1({**meta["subject"], "source_file": key["file_id"]})
Session.insert1({**meta["session"], "source_file": key["file_id"]})

self.insert1({**key, "n_entities": 2})


def parse(file_id):
"""Stand-in for a real parser."""
return {
"subject": {"subject_id": 100 + file_id, "species": "mouse"},
"session": {"session_id": 200 + file_id, "session_date": "2026-09-30"},
}


if __name__ == "__main__":
RecordingFile.insert1({"file_id": 1, "path": "/mnt/raw/session-001.nwb"})
Ingest.populate()

# The hidden attribute is excluded from the heading, so read it explicitly.
for row in (Subject & "subject_id = 101").proj("_prov").to_dicts():
print(row["_prov"])
# {'time': '2026-09-30T14:22:05.481203+00:00',
# 'agent': {'user': ..., 'host': ..., 'database_name': ...},
# 'source': {'system': 'AcquisitionShare', 'root': '/mnt/raw'},
# 'context': {'table': '`fan_out_provenance_demo`.`_ingest`',
# 'key': {'file_id': 1}}}

# The audit question the pattern used to leave unanswerable.
print("rows with no recorded origin:", len(Subject & "_prov IS NULL"))

schema.drop()
2 changes: 2 additions & 0 deletions mkdocs.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -79,6 +79,7 @@ nav:
- Deploy to Production: how-to/deploy-production.md
- Data Operations:
- Insert Data: how-to/insert-data.md
- Record Data Origin: how-to/record-data-origin.md
- Query Data: how-to/query-data.md
- Fetch Results: how-to/fetch-results.md
- Delete Data: how-to/delete-data.md
Expand Down Expand Up @@ -132,6 +133,7 @@ nav:
- AutoPopulate: reference/specs/autopopulate.md
- Upstream Trace: reference/specs/trace.md
- Job Metadata: reference/specs/job-metadata.md
- Entry Provenance: reference/specs/boundary-provenance.md
- Object Store Configuration: reference/specs/object-store-configuration.md
- Deployment:
- Deployment Operations: reference/specs/deploy-operations.md
Expand Down
12 changes: 8 additions & 4 deletions src/explanation/comparison-to-provenance-systems.md
Original file line number Diff line number Diff line change
Expand Up @@ -53,11 +53,15 @@ within-pipeline derivation and external-origin provenance are different
questions, and a complete record often wants both.

Inside DataJoint, the origin of externally-sourced data is recorded at the
pipeline's entry-point tables — a Manual insert or an Imported `make()` records
the source identity alongside the data, exactly as at any manual data-entry
point. A single ingestion step may populate several such tables that carry no
pipeline's entry-point tables. Since 2.3.4 that record has a standard shape: a
hidden `_prov` attribute on every Manual table, written by the framework from
deployment configuration and from the executing context rather than by the
author — see
[Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md).
A single ingestion step may populate several such tables that carry no
foreign-key dependency on the loader (the
[fan-out ingestion pattern](fan-out-ingestion.md)), each recording its own origin.
[fan-out ingestion pattern](fan-out-ingestion.md)); rows written from inside that
loader record it, which is the link the absent foreign key would have carried.

At the boundary, the two integrate **in both directions** — when explicitly
configured:
Expand Down
4 changes: 4 additions & 0 deletions src/explanation/computation-model.md
Original file line number Diff line number Diff line change
Expand Up @@ -228,6 +228,10 @@ This adds to computed tables:
- `_job_duration` — How long it took
- `_job_version` — Code version (if configured)

They are hidden: filtered out of the heading, out of `to_dicts()`, and out of
join matching. See [Hidden Job Metadata](../reference/specs/job-metadata.md) for
how to query them.

## The Three-Part Make Model

For long-running computations (hours or days), holding a database transaction
Expand Down
37 changes: 24 additions & 13 deletions src/explanation/fan-out-ingestion.md
Original file line number Diff line number Diff line change
Expand Up @@ -61,21 +61,32 @@ carry whichever loader existed when it was created. The pattern avoids that by
*declining* the FK on purpose, rather than letting an undeclared dependency slip
in unnoticed.

## The responsibility it carries: record where the data came from
## The origin is recorded for you, and modelled by you

Because the foreign-key link to the source is absent, the traceability it would
have provided must be supplied another way. **Each table the pattern populates is
responsible for recording its own origin** — the source identity the row was
derived from (file path, checksum, instrument session, operator, timestamp,
external record id). This is the same responsibility every `Manual` and
`Imported` entry-point table already carries: data entering the pipeline from
outside must record where it came from, because the pipeline's own structure
cannot vouch for it.

Recording that origin at the point of entry is all DataJoint asks. Formalizing
and standardizing it beyond that — retention, audit trails, cross-system
exchange — is left to the provenance and governance systems a pipeline
interoperates with; see [Comparison to Provenance Systems](comparison-to-provenance-systems.md).
have provided has to come from somewhere else. Two things supply it, and they do
different jobs.

**DataJoint records the origin automatically.** Every row written into a `Manual`
table carries a hidden `_prov` attribute, and a row written from inside an
ingesting `make()` records the ingesting table and its key — exactly the link the
missing foreign key would have carried. Nothing in the `make()` body asks for
this, and nothing can forge it: the attribute is framework-owned and no insert
can set it. What it records beyond that comes from deployment configuration —
the external system, the connecting user, the time, the code version. See
[Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md).

**You model the link the pipeline itself needs to query.** Hidden attributes are
deliberately excluded from query composition, so `_prov` cannot be joined or
restricted on the way an ordinary attribute can. Where downstream code has to
follow the row back to its source — and in the example above it does — keep the
`source_file` column. `_prov` is the audit record; the modelled column is the
domain link.

Beyond recording the origin at the point of entry, the rest — retention, audit
trails, cross-system exchange — is left to the provenance and governance systems
a pipeline interoperates with; see
[Comparison to Provenance Systems](comparison-to-provenance-systems.md).

## When to use it

Expand Down
2 changes: 2 additions & 0 deletions src/how-to/monitor-progress.md
Original file line number Diff line number Diff line change
Expand Up @@ -108,6 +108,8 @@ This adds hidden attributes to computed tables:
- `_job_duration` — How long it took
- `_job_version` — Code version (if configured)

Restricting on one works as a condition string — `SessionAnalysis & "_job_duration > 3600"` — but reading the values back needs SQL until 2.4. See [Querying and Fetching](../reference/specs/job-metadata.md#querying-and-fetching).

## Simple Progress Script

```python
Expand Down
137 changes: 137 additions & 0 deletions src/how-to/record-data-origin.md
Original file line number Diff line number Diff line change
@@ -0,0 +1,137 @@
# Record Data Origin

Data entering a pipeline from outside carries no dependency that says where it came from. Turn capture on and DataJoint records that origin for you on every Manual table, from configuration rather than from your insert code.

!!! version-added "New in 2.3.4"

## Turn capture on

Capture is off by default, so a table declared without it carries no column and records nothing. Enable it where you set credentials and stores, before the tables are declared:

```bash
export DJ_PROVENANCE_CAPTURE=true
```

Turning it off again does not remove the column from tables that already have it, and does not stop those tables from recording.

## Configure the source

Name the external system this process draws from. Set it where you set credentials and stores — not in pipeline code:

```bash
export DJ_PROVENANCE_SOURCE='{"system": "PyRat", "endpoint": "https://pyrat.example.org/api/v2"}'
```

or in `datajoint.json`:

```json
{
"provenance": {
"source": {"system": "PyRat", "endpoint": "https://pyrat.example.org/api/v2"}
}
}
```

Every row this process inserts into a Manual table now records that source, along with the connecting user and host, the insert time, and the code version.

Nothing else is required. There is no argument to pass and no field to remember:

```python
Subject.insert1({"subject_id": 1, "species": "mouse"})
```

Even when the source offers little — a nightly sync against a colony-management API — recording "received from PyRat at 02:15" beats recording nothing.

## Check that rows are carrying an origin

The attribute is hidden, so it does not appear in `to_dicts()` or in a join. Query it directly:

```python
# Rows with no recorded origin
Subject & "_prov IS NULL"

# How many, out of how many
len(Subject & "_prov IS NULL"), len(Subject)
```

Write the condition as a **string**. The mapping form returns every row here: it
ignores attributes it cannot match — deliberately, so that `Session & key` works
when `key` carries attributes from a more detailed table — and a hidden
attribute is invisible to that matching
([#1561](https://github.com/datajoint/datajoint-python/issues/1561)):

```python
# MySQL
Subject & "JSON_VALUE(_prov, '$.source.system') = 'PyRat'"

# PostgreSQL
Subject & "jsonb_extract_path_text(_prov, 'source', 'system') = 'PyRat'"

# Subject & {"_prov.system": "PyRat"} <- returns everything; do not use
```

`_prov IS NULL` and `_prov IS NOT NULL` are the same on both backends. Filtering
on a field inside the JSON is not — **because `_prov` is hidden**, not because
JSON paths are hard. On an ordinary JSON attribute `{"data.system": "PyRat"}` is
portable and DataJoint translates it per backend; that route is closed here only
because the mapping form cannot reach a hidden attribute.

## Read the record back

Until 2.4 this needs SQL. `to_arrays("_prov")` and `proj("_prov")` both raise,
because a hidden attribute cannot be named through the query API
([#1562](https://github.com/datajoint/datajoint-python/issues/1562) adds a
supported accessor):

```python
rows = Subject.connection.query(
f"SELECT subject_id, _prov FROM {Subject.full_table_name}"
).fetchall()
```

On MySQL the value comes back as a JSON string and needs `json.loads`; on
PostgreSQL psycopg2 returns a dict already.

## Add the column to existing tables

Tables declared before 2.3.4 — or while capture was off — have no column, and inserts into them record nothing without complaining. Add the slot:

```python
from datajoint.deploy import add_prov_column

# See what would change
add_prov_column(schema, dry_run=True)["ddl"]

# Apply it
add_prov_column(schema, dry_run=False)
```

Safe to re-run: a table that already has the column is reported and left alone. Rows already present keep `NULL` — provenance is recorded when a row is inserted and is never reconstructed afterwards.

## What you cannot do, and what to do instead

**You cannot write `_prov` yourself.** Passing it in a row raises an error.

That is deliberate. A field the operator can set is weaker evidence than one the system sets, which is the whole point for an audit. It also means the record cannot be half-filled by inconsistent discipline across a team.

**When you want to record something specific to a row** — which file a value came from, which LIMS record, which operator — model it as an ordinary attribute:

```python
@schema
class Subject(dj.Manual):
definition = """
subject_id : int32
---
species : varchar(64)
lims_record : varchar(64) # the external record this row was created from
"""
```

A modeled column is visible, queryable, and joinable; `_prov` is none of those, by design. Use `_prov` as the audit record and a modeled column as the domain link. Both can describe the same arrival.

## See Also

- [Extrinsic Provenance at Entry Tables](../reference/specs/boundary-provenance.md) — the specification
- [Insert Data](insert-data.md) — inserting into Manual tables
- [Fan-Out Ingestion](../explanation/fan-out-ingestion.md) — one loader writing into several entry-point tables
- [Configuration](../reference/configuration.md) — where settings come from
17 changes: 17 additions & 0 deletions src/reference/configuration.md
Original file line number Diff line number Diff line change
Expand Up @@ -163,6 +163,23 @@ If table lacks partition attributes, it follows normal path structure.
| `jobs.add_job_metadata` | `False` | Add hidden metadata to computed tables |
| `jobs.allow_new_pk_fields_in_computed_tables` | `False` | Allow non-FK primary key fields |

## Provenance Settings

| Setting | Environment | Default | Description |
| --------- | ------------- | --------- | ------------- |
| `provenance.capture` | `DJ_PROVENANCE_CAPTURE` | `False` | Declare the hidden `_prov` attribute on Manual tables and fill it on insert *(new in 2.3.4)* |
| `provenance.source` | `DJ_PROVENANCE_SOURCE` | `{}` | External source identity recorded on every row this process enters *(new in 2.3.4)* |

`provenance.source` is a JSON object naming the system this process draws from,
set per deployment rather than in pipeline code:

```bash
export DJ_PROVENANCE_SOURCE='{"system": "PyRat", "endpoint": "https://pyrat.example.org/api/v2"}'
```

See [Extrinsic Provenance at Entry Tables](specs/boundary-provenance.md) for what is
recorded and [Record Data Origin](../how-to/record-data-origin.md) for the task-oriented guide.

## Display Settings

| Setting | Environment | Default | Description |
Expand Down
17 changes: 12 additions & 5 deletions src/reference/specs/autopopulate.md
Original file line number Diff line number Diff line change
Expand Up @@ -988,17 +988,24 @@ When `config['jobs.add_job_metadata'] = True`, auto-populated tables receive hid
| Column | Type | Description |
|--------|------|-------------|
| `_job_start_time` | `datetime(3)` | When computation began |
| `_job_duration` | `float64` | Duration in seconds |
| `_job_duration` | `float32` | Duration in seconds |
| `_job_version` | `varchar(64)` | Code version |

```python
# Fetch with job metadata
Analysis().to_arrays('result', '_job_duration')

# Query slow computations
# Query slow computations -- a condition string reaches the column
slow = Analysis & '_job_duration > 3600'

# Reading the values back requires SQL until 2.4
rows = Analysis.connection.query(
f"SELECT _job_start_time, _job_duration FROM {Analysis.full_table_name}"
).fetchall()
```

`to_arrays('_job_duration')` and `proj('_job_duration')` raise: the heading
excludes hidden names, so they cannot be addressed through the query API. See
[Hidden Job Metadata](job-metadata.md#querying-and-fetching) for the full
account and for what 2.4 adds.

---

## 15. Migration from Legacy DataJoint
Expand Down
Loading