Skip to content

Add a hidden _prov attribute for extrinsic provenance at pipeline boundaries #1547

Description

@dimitri-yatsenko

The gap

Intrinsic provenance is structural: a Compute table's row cannot exist unless its declared upstream exists and is correct, so the foreign-key graph is the lineage. Nothing needs to be recorded for that to hold.

At the boundary, the structure runs out. Rows arrive from outside the pipeline — an Entry table filled by a person or a feed, an Ingest table's make() reading a file the workflow does not track, or a fan-out write into Entry tables that carry no foreign key back to the writer. For those rows the framework has no way to say where they came from, and today each pipeline invents its own: a source_file column here, a notes varchar there, an ingestion log somewhere else, or nothing at all.

That is the one place where DataJoint's provenance story depends entirely on the workflow author's diligence, with no shape to conform to.

The precedent this should follow

The mechanism already exists in the codebase. config.jobs.add_job_metadata adds hidden attributes to Computed/Imported tables at declaration:

_job_start_time  datetime(3)
_job_duration    float
_job_version     varchar(64)

(src/datajoint/adapters/base.py, and per-adapter in mysql.py / the PostgreSQL path.) These are per-row, hidden from heading, written by populate, and durable on the table itself rather than in the job queue. That is the right shape and the right place — it just covers the automated side, where provenance is already intrinsic, and not the boundary, where it is not.

Proposal

A hidden _prov attribute, JSON-typed, on tables that take rows from outside the pipeline.

_prov  json  DEFAULT NULL   # extrinsic provenance for a row that entered from outside
  • Where: available on Entry and Ingest tables. A Compute table has no use for it, since its provenance is entailed.
  • Who writes it: the workflow author, at the point of entry — insert on an Entry table, or a fan-out write from inside a make(). The framework provides the slot and the shape; it cannot infer content it did not produce.
  • What it holds: at minimum the agent, the external source and record identifier, the time, and by what means. A conventional key set matters more than a rigid schema — the point is that two pipelines answering "where did this row come from" answer it in the same shape.
  • Master carries it; parts inherit. Consistent with how the master/part relationship works elsewhere.
  • Fan-out: inside make(), the intrinsic record of the ingesting table (its key, its _job_version) is available to propagate into each extrinsic destination, which is what makes a fanned-out row traceable without a foreign key.

Even when the source offers nothing useful — a nightly sync against a colony-management API — the author still records "received from PyRat at 02:15". A slot with a weak value beats no slot.

Why in the framework rather than left to each pipeline

Three things follow from having one shape:

  1. Export becomes mechanical. W3C PROV wants wasAttributedTo and wasDerivedFrom on exactly these rows; OpenLineage wants the same content run-centric. With per-pipeline conventions, every export is bespoke.
  2. The boundary becomes inspectable. "Which Entry rows have no recorded origin" turns into a query rather than an audit.
  3. ALCOA+ attributability lands where it belongs. Deployments that must answer attributable and contemporaneous for externally-sourced data currently have nowhere standard to put the answer.

Open questions

  • Default on or off? add_job_metadata defaults to False, and tables declared without it never get the columns — a migration edge worth not repeating. A _prov slot that is absent on most tables is a slot nobody codes against.
  • Validation. Enforce a minimal key set at insert, or accept any JSON and let deployments constrain it? Leaning toward the latter in the framework, with strictness as a deployment concern.
  • Interaction with allow_direct_insert. A direct insert into an Ingest table is already a modeling smell (datajoint-docs#267); should it require _prov?
  • Naming. _prov is short and matches the _job_* convention. _source or _origin would read more plainly to someone who has not met the term.

Related: datajoint-docs#267 (tier names on the entry/ingest axis), and the fan-out ingestion explanation, which currently tells authors to record source identity without giving them a place to record it.

Activity

  1. added this to the v2.3.4 milestone on Sep 30, 2026
  2. dimitri-yatsenko commented on Sep 30, 2026

    @dimitri-yatsenko
    MemberAuthor

    Settled design for 2.3.4

    Two amendments to the proposal above, and the fill mechanism settled. Scope narrows to Entry tables only, and the attribute becomes framework-owned: nothing an author writes, at declaration or at insert.

    Storage: one hidden JSON column on the table

    _prov  json  DEFAULT NULL
    

    Not a side table. An append-only ingestion log would need its own tier, its own cascade semantics, and an orphan policy. "Where did this row come from" is an attribute of the row, and re-ingesting a key is an update, not a second origin. Survival after delete belongs to the garbage-collection work (#1445, #1478), not here. If the Platform later wants an accruing ingestion-event log for re-syncs and corrections, it builds on this shape rather than replacing it.

    Not a set of typed _prov_* columns. The key set is deliberately unsettled — this issue asks for "at minimum the agent, the external source and record identifier, the time, and by what means," and a conventional key set matters more than a rigid schema. Typed columns freeze that set at declaration and turn every later addition into an ALTER across every Entry table in every deployment. JSON makes it a configuration change. Field filtering stays portable: condition.py already translates JSON path restrictions to json_value() on MySQL and jsonb_extract_path_text() on PostgreSQL.

    The _job_* precedent does not transfer here. Those are three fixed framework-owned scalars that will never grow; provenance content is open-ended by design.

    Content comes from three sources, none of them the call site

    Source Supplies
    Settings — config.provenance.* the deployment constant: which external system, which process
    Ambient connection state conn_info user, host, and database; insert timestamp; _get_job_version(config)
    Ambient execution state the executing table and its key, when the insert runs inside a make()

    insert's signature does not change and there is no prov= keyword.

    The third source is what makes fan-out ingestion traceable without a foreign key: rows written into Entry tables from inside an ingesting make() record the ingesting table and key automatically, because the framework knows both at that moment.

    The invariant: an author cannot write _prov

    A field the operator can set is weaker evidence than one the system sets. Framework ownership is what makes the record worth trusting for ALCOA+ attributability, and it removes the adoption risk that an author-supplied field carries — there is nothing left for a pipeline to neglect.

    Anything an author wants to record deliberately belongs in the model as a visible attribute, where pipeline code can restrict and join on it. Hidden attributes are excluded from query composition by design (heading.py:267, and the binary-operator handling documented in the job-metadata spec), so _prov is the audit record and a modeled column is the domain link. The two do different jobs.

    Entry only; Ingest does not need it

    Three reasons compound.

    The agent, time, and version half is already covered on Ingest. config.jobs.add_job_metadata adds _job_start_time, _job_duration, and _job_version to Computed and Imported tables. A settings-driven _prov on Ingest would record the same facts twice.

    The half not covered is the half settings cannot supply. An Ingest table's extrinsic provenance wants the specific external resource its make() read — which file, which endpoint, which instrument session — known per row, inside the make() body. Settings supplies a process-level constant, so Ingest would gain the duplicated half and miss the useful one.

    A well-modeled pipeline registers the external source as an Entry row. In the fan-out example, RecordingFile(dj.Manual) holds the path and the ingesting table declares a foreign key to it, which makes that table's provenance intrinsic. The external world enters at the Entry table.

    Where an Ingest make() reads something not registered as an Entry row, that is the modeling problem described in datajoint/datajoint-docs#267. The fix is to register that source as an Entry table, which then carries _prov and restores the declared dependency. Giving Ingest a provenance slot instead would hand the modeling problem somewhere to hide, and record duplicated facts while calling the boundary covered.

    The remaining per-row question — which external resource a make() actually read — is consumed-input capture. Different mechanism, different capability.

    On by default

    config.jobs.add_job_metadata defaults to False, which is why migrate.add_job_metadata_columns exists. Repeating that leaves no consumer able to assume the column exists, and turns "which Entry rows have no recorded origin" into a question conditional on each table's declaration-time configuration rather than a query.

    Capture defaults on for Entry tables, with a configuration flag to disable it. Because the slot fills itself, an author who does nothing still gets agent, time, and version.

    Stating the consequence plainly: an Entry table declared under 2.3.4 differs in DDL from one declared under 2.3.3. The difference is additive and hidden.

    Where the boundary sits

    The library fixes the shape and reads configuration. The Platform sets that configuration per project, through the channel it already uses for stores and credentials — DJ_PROVENANCE_*, the configuration file, or the secrets directory. Enforcement stays commercial and the library holds no policy.

    Implementation

    File Change
    settings.py ProvenanceSettings, env_prefix="DJ_PROVENANCE_", exposed as config.provenance
    declare.py add _prov to Entry tables at declaration — mirrors the add_job_metadata branch at line 516
    adapters/{base,mysql,postgres}.py column DDL — mirrors the _job_* blocks
    table.py assemble and write the value on the insert path
    migrate.py add_prov_column retrofit — mirrors add_job_metadata_columns

    Still open

    • Configuration change mid-process shifts what later rows record, silently. This wants a note in the spec rather than a mechanism.
    • Range queries on capture time go through JSON extraction. If that becomes hot, a deployment adds a generated column and indexes it, with no library change.
    • allow_direct_insert interaction stays out of scope for 2.3.4. Requiring _prov on a direct insert is enforcement, which sits with the Platform.
    • Naming. _prov carries, since the author never types it and provenance is the term the rest of the model uses.

    Documentation

    datajoint/datajoint-docs#284: a new specification, the config.provenance reference, a tutorial extension, a worked example, and a rewrite of the fan-out ingestion page, whose "responsibility it carries" section currently assigns the recording to the author.

  3. dimitri-yatsenko commented on Sep 30, 2026

    @dimitri-yatsenko
    MemberAuthor

    Implementation opened as #1555, with the design above implemented as settled. Documentation is tracked in datajoint/datajoint-docs#284 — the reference spec lands there as reference/specs/boundary-provenance.md, alongside a how-to, the config.provenance reference, and a rewrite of the fan-out ingestion page, whose "responsibility it carries" section currently assigns the recording to the author.

    Two departures from the text above, both noted in the PR:

    • The retrofit helper is deploy.add_prov_column, not a migrate function. datajoint.migrate is deprecated for removal in 2.4/2.5, but every existing deployment needs this retrofit and will still need it afterwards whenever capture has been off. The helper is idempotent, which is deploy's stated contract.
    • _prov_* typed columns became a single _prov JSON column, per the discussion above: the key set is deliberately unsettled, and typed columns would freeze it at declaration and turn every later addition into an ALTER across every Entry table in every deployment.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    No projects

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions