Repository navigation
Add a hidden _prov attribute for extrinsic provenance at pipeline boundaries #1547
Description
Activity
Settled design for 2.3.4
Two amendments to the proposal above, and the fill mechanism settled. Scope narrows to Entry tables only, and the attribute becomes framework-owned: nothing an author writes, at declaration or at insert.
Storage: one hidden JSON column on the table
_prov json DEFAULT NULLNot a side table. An append-only ingestion log would need its own tier, its own cascade semantics, and an orphan policy. "Where did this row come from" is an attribute of the row, and re-ingesting a key is an update, not a second origin. Survival after delete belongs to the garbage-collection work (#1445, #1478), not here. If the Platform later wants an accruing ingestion-event log for re-syncs and corrections, it builds on this shape rather than replacing it.
Not a set of typed
_prov_*columns. The key set is deliberately unsettled — this issue asks for "at minimum the agent, the external source and record identifier, the time, and by what means," and a conventional key set matters more than a rigid schema. Typed columns freeze that set at declaration and turn every later addition into anALTERacross every Entry table in every deployment. JSON makes it a configuration change. Field filtering stays portable:condition.pyalready translates JSON path restrictions tojson_value()on MySQL andjsonb_extract_path_text()on PostgreSQL.The
_job_*precedent does not transfer here. Those are three fixed framework-owned scalars that will never grow; provenance content is open-ended by design.Content comes from three sources, none of them the call site
Source Supplies Settings — config.provenance.*the deployment constant: which external system, which process Ambient connection state conn_infouser, host, and database; insert timestamp;_get_job_version(config)Ambient execution state the executing table and its key, when the insert runs inside a make()insert's signature does not change and there is noprov=keyword.The third source is what makes fan-out ingestion traceable without a foreign key: rows written into Entry tables from inside an ingesting
make()record the ingesting table and key automatically, because the framework knows both at that moment.The invariant: an author cannot write
_provA field the operator can set is weaker evidence than one the system sets. Framework ownership is what makes the record worth trusting for ALCOA+ attributability, and it removes the adoption risk that an author-supplied field carries — there is nothing left for a pipeline to neglect.
Anything an author wants to record deliberately belongs in the model as a visible attribute, where pipeline code can restrict and join on it. Hidden attributes are excluded from query composition by design (
heading.py:267, and the binary-operator handling documented in the job-metadata spec), so_provis the audit record and a modeled column is the domain link. The two do different jobs.Entry only; Ingest does not need it
Three reasons compound.
The agent, time, and version half is already covered on Ingest.
config.jobs.add_job_metadataadds_job_start_time,_job_duration, and_job_versionto Computed and Imported tables. A settings-driven_provon Ingest would record the same facts twice.The half not covered is the half settings cannot supply. An Ingest table's extrinsic provenance wants the specific external resource its
make()read — which file, which endpoint, which instrument session — known per row, inside themake()body. Settings supplies a process-level constant, so Ingest would gain the duplicated half and miss the useful one.A well-modeled pipeline registers the external source as an Entry row. In the fan-out example,
RecordingFile(dj.Manual)holds the path and the ingesting table declares a foreign key to it, which makes that table's provenance intrinsic. The external world enters at the Entry table.Where an Ingest
make()reads something not registered as an Entry row, that is the modeling problem described in datajoint/datajoint-docs#267. The fix is to register that source as an Entry table, which then carries_provand restores the declared dependency. Giving Ingest a provenance slot instead would hand the modeling problem somewhere to hide, and record duplicated facts while calling the boundary covered.The remaining per-row question — which external resource a
make()actually read — is consumed-input capture. Different mechanism, different capability.On by default
config.jobs.add_job_metadatadefaults toFalse, which is whymigrate.add_job_metadata_columnsexists. Repeating that leaves no consumer able to assume the column exists, and turns "which Entry rows have no recorded origin" into a question conditional on each table's declaration-time configuration rather than a query.Capture defaults on for Entry tables, with a configuration flag to disable it. Because the slot fills itself, an author who does nothing still gets agent, time, and version.
Stating the consequence plainly: an Entry table declared under 2.3.4 differs in DDL from one declared under 2.3.3. The difference is additive and hidden.
Where the boundary sits
The library fixes the shape and reads configuration. The Platform sets that configuration per project, through the channel it already uses for stores and credentials —
DJ_PROVENANCE_*, the configuration file, or the secrets directory. Enforcement stays commercial and the library holds no policy.Implementation
File Change settings.pyProvenanceSettings,env_prefix="DJ_PROVENANCE_", exposed asconfig.provenancedeclare.pyadd _provto Entry tables at declaration — mirrors theadd_job_metadatabranch at line 516adapters/{base,mysql,postgres}.pycolumn DDL — mirrors the _job_*blockstable.pyassemble and write the value on the insert path migrate.pyadd_prov_columnretrofit — mirrorsadd_job_metadata_columnsStill open
- Configuration change mid-process shifts what later rows record, silently. This wants a note in the spec rather than a mechanism.
- Range queries on capture time go through JSON extraction. If that becomes hot, a deployment adds a generated column and indexes it, with no library change.
allow_direct_insertinteraction stays out of scope for 2.3.4. Requiring_provon a direct insert is enforcement, which sits with the Platform.- Naming.
_provcarries, since the author never types it and provenance is the term the rest of the model uses.
Documentation
datajoint/datajoint-docs#284: a new specification, the
config.provenancereference, a tutorial extension, a worked example, and a rewrite of the fan-out ingestion page, whose "responsibility it carries" section currently assigns the recording to the author.Implementation opened as #1555, with the design above implemented as settled. Documentation is tracked in datajoint/datajoint-docs#284 — the reference spec lands there as
reference/specs/boundary-provenance.md, alongside a how-to, theconfig.provenancereference, and a rewrite of the fan-out ingestion page, whose "responsibility it carries" section currently assigns the recording to the author.Two departures from the text above, both noted in the PR:
- The retrofit helper is
deploy.add_prov_column, not amigratefunction.datajoint.migrateis deprecated for removal in 2.4/2.5, but every existing deployment needs this retrofit and will still need it afterwards whenever capture has been off. The helper is idempotent, which isdeploy's stated contract. _prov_*typed columns became a single_provJSON column, per the discussion above: the key set is deliberately unsettled, and typed columns would freeze it at declaration and turn every later addition into anALTERacross every Entry table in every deployment.
- The retrofit helper is
The gap
Intrinsic provenance is structural: a Compute table's row cannot exist unless its declared upstream exists and is correct, so the foreign-key graph is the lineage. Nothing needs to be recorded for that to hold.
At the boundary, the structure runs out. Rows arrive from outside the pipeline — an Entry table filled by a person or a feed, an Ingest table's
make()reading a file the workflow does not track, or a fan-out write into Entry tables that carry no foreign key back to the writer. For those rows the framework has no way to say where they came from, and today each pipeline invents its own: asource_filecolumn here, anotesvarchar there, an ingestion log somewhere else, or nothing at all.That is the one place where DataJoint's provenance story depends entirely on the workflow author's diligence, with no shape to conform to.
The precedent this should follow
The mechanism already exists in the codebase.
config.jobs.add_job_metadataadds hidden attributes to Computed/Imported tables at declaration:(
src/datajoint/adapters/base.py, and per-adapter inmysql.py/ the PostgreSQL path.) These are per-row, hidden fromheading, written bypopulate, and durable on the table itself rather than in the job queue. That is the right shape and the right place — it just covers the automated side, where provenance is already intrinsic, and not the boundary, where it is not.Proposal
A hidden
_provattribute, JSON-typed, on tables that take rows from outside the pipeline.inserton an Entry table, or a fan-out write from inside amake(). The framework provides the slot and the shape; it cannot infer content it did not produce.make(), the intrinsic record of the ingesting table (its key, its_job_version) is available to propagate into each extrinsic destination, which is what makes a fanned-out row traceable without a foreign key.Even when the source offers nothing useful — a nightly sync against a colony-management API — the author still records "received from PyRat at 02:15". A slot with a weak value beats no slot.
Why in the framework rather than left to each pipeline
Three things follow from having one shape:
wasAttributedToandwasDerivedFromon exactly these rows; OpenLineage wants the same content run-centric. With per-pipeline conventions, every export is bespoke.Open questions
add_job_metadatadefaults toFalse, and tables declared without it never get the columns — a migration edge worth not repeating. A_provslot that is absent on most tables is a slot nobody codes against.allow_direct_insert. A direct insert into an Ingest table is already a modeling smell (datajoint-docs#267); should it require_prov?_provis short and matches the_job_*convention._sourceor_originwould read more plainly to someone who has not met the term.Related: datajoint-docs#267 (tier names on the entry/ingest axis), and the fan-out ingestion explanation, which currently tells authors to record source identity without giving them a place to record it.