diff --git a/src/about/versioning.md b/src/about/versioning.md index 365618c2..6a49c115 100644 --- a/src/about/versioning.md +++ b/src/about/versioning.md @@ -58,7 +58,7 @@ Version admonitions are used for features introduced **after 2.0** (i.e., versio !!! version-deprecated "Deprecated in 2.1, removed in 3.0" - The `allow_direct_insert` parameter is deprecated. Use `dj.config['safemode']` instead. + The `legacy_option` parameter is deprecated. Use `replacement_option` instead. **Note:** Features deprecated at the 2.0 baseline (coming from pre-2.0) are documented in the [Migration Guide](../how-to/migrate-to-v20.md) rather than with admonitions, since this documentation assumes 2.0 as the baseline. diff --git a/src/explanation/relational-workflow-model.md b/src/explanation/relational-workflow-model.md index fac1d7f7..72bc2964 100644 --- a/src/explanation/relational-workflow-model.md +++ b/src/explanation/relational-workflow-model.md @@ -145,17 +145,44 @@ work in practice. ### Workflow steps and table tiers -Tables are classified into tiers by data-entry mode: +Tables are classified into tiers by what puts rows in them. -| Tier | Role | `make()` | -|------|------|----------| -| **Manual** | Rows entered at runtime from outside the pipeline (people, forms, instruments, imports) | No | -| **Lookup** | Reference rows defined in the schema itself via `contents` | No | -| **Imported** | Reach out to data sources outside DataJoint (instruments, ELNs, external databases) | Yes | -| **Computed** | Derive their contents entirely from upstream DataJoint tables | Yes | +| Tier | Rows come from | What puts them there | `make()` | +|------|----------------|----------------------|----------| +| **Lookup** | The committed schema | The table's own `contents`, versioned with the code | No | +| **Manual** | Outside the pipeline | A writer outside the table — a person, an instrument, an entry script | No | +| **Imported** | Outside the pipeline | The table itself, fetching through `make()` | Yes | +| **Computed** | Other DataJoint tables | The table itself, deriving through `make()` | Yes | -Imported and Computed tables define computations via `make()` methods. The -`make()` method specifies how each entity is derived — declared within the +The table shows **Manual** and **Imported** drawing on the same origin. What +separates them is who initiates the write. + +A Manual table is written by an external process, on that process's schedule. +An Imported table is filled automatically: `populate()` works through the keys +its parents already hold and calls `make()` for each one still missing. + +Their primary keys follow from that. `populate()` has to know which entity it +is working on before `make()` runs, so every attribute of an Imported table's +primary key arrives through a foreign key. A Manual table carries no such +constraint — it may sit at the head of the pipeline with no parent at all, and +may introduce primary-key attributes of its own. That is what makes it the +place a new entity enters. + +!!! version-added "New in 2.3.4" + + Three tiers gain a second name: **`dj.Entry`** for `dj.Manual`, + **`dj.Ingest`** for `dj.Imported`, and **`dj.Compute`** for `dj.Computed`. + Each pair is one class, so either name declares the same table, and both + names are permanent. `dj.Lookup` and `dj.Part` are unchanged. + + These pages use the original names. The new ones become primary in 2.4 + ([datajoint-python#1546](https://github.com/datajoint/datajoint-python/issues/1546)). + +`Part` is absent because it is not a tier of its own. A part table fills a +structural role: it inherits its master's tier and is written in the same +transaction. Any tier can serve as a master. + +The `make()` method specifies how each entity is derived — declared within the table definition, not in an external workflow file. #### Manual vs. Lookup diff --git a/src/how-to/define-tables.md b/src/how-to/define-tables.md index 7fb5f850..a0833adc 100644 --- a/src/how-to/define-tables.md +++ b/src/how-to/define-tables.md @@ -30,7 +30,7 @@ class MyTable(dj.Manual): | Type | Base Class | Purpose | |------|------------|---------| -| Manual | `dj.Manual` | Data inserted directly from outside the pipeline (forms, instruments, ingestion scripts) | +| Manual | `dj.Manual` | Data inserted directly from outside the pipeline (forms, instruments, entry scripts) | | Lookup | `dj.Lookup` | Reference data defined in the schema via `contents` | | Imported | `dj.Imported` | Populated by `make()` from an external source | | Computed | `dj.Computed` | Populated by `make()` from other tables | @@ -266,8 +266,8 @@ data management?** process, not a runtime insert. - Use **`dj.Manual`** when the rows are **populated at runtime** and their quality is guaranteed by the **data-management process** — validation, curation, and - access control at ingest, not code review. Subjects, sessions, samples, or - anything typed into a form, ingested from a file, or read from an instrument. + access control at entry, not code review. Subjects, sessions, samples, or + anything typed into a form, loaded from a file, or read from an instrument. This resolves the case that "where a row comes from" leaves ambiguous: a controlled vocabulary that is **populated at runtime** — gene symbols loaded from an external diff --git a/src/how-to/insert-data.md b/src/how-to/insert-data.md index fa9cff1b..221c7274 100644 --- a/src/how-to/insert-data.md +++ b/src/how-to/insert-data.md @@ -59,7 +59,7 @@ Subject.insert(rows, ignore_extra_fields=True) The sections above target **Manual** tables — the tier whose rows are inserted directly, from outside the DataJoint pipeline (whether entered by hand through a -form or GUI, or loaded by an automated ingestion tool). Other tiers get their +form or GUI, or loaded by an automated entry script). Other tiers get their rows a different way, and inserting into them directly breaks reproducibility: - **Computed and Imported tables** — their rows are produced only by `make()` diff --git a/src/how-to/installation.md b/src/how-to/installation.md index b3a78aa5..8ba8d0ab 100644 --- a/src/how-to/installation.md +++ b/src/how-to/installation.md @@ -159,7 +159,7 @@ Session.insert1({'subject_id': 1, 'session_idx': 1, 'session_date': '2026-01-06' SessionAnalysis.populate() ``` -`Subject` and `Session` are entered by hand; `SessionAnalysis` derives from `Session` and fills +`Subject` and `Session` are written from outside the pipeline; `SessionAnalysis` derives from `Session` and fills itself when you call `populate()`. That dependency — declared with `->` — is the whole of the [Relational Workflow Model](../explanation/relational-workflow-model.md) in miniature. diff --git a/src/reference/specs/table-declaration.md b/src/reference/specs/table-declaration.md index 6bbb9833..0d994c14 100644 --- a/src/reference/specs/table-declaration.md +++ b/src/reference/specs/table-declaration.md @@ -23,7 +23,7 @@ class TableName(dj.Manual): | Tier | Base Class | Table Prefix | Purpose | |------|------------|--------------|---------| -| Manual | `dj.Manual` | (none) | Data inserted at runtime from outside the pipeline (users, instruments, ingestion scripts); quality assured by the data-management process | +| Manual | `dj.Manual` | (none) | Data inserted at runtime from outside the pipeline (users, instruments, entry scripts); quality assured by the data-management process | | Lookup | `dj.Lookup` | `#` | Reference data defined in the schema via `contents`; quality assured by code review | | Imported | `dj.Imported` | `_` | Populated by `make()` from an external source | | Computed | `dj.Computed` | `__` | Derived from other tables | diff --git a/src/tutorials/advanced/sql-comparison.ipynb b/src/tutorials/advanced/sql-comparison.ipynb index 68dbd121..a2fc5fcd 100644 --- a/src/tutorials/advanced/sql-comparison.ipynb +++ b/src/tutorials/advanced/sql-comparison.ipynb @@ -2067,7 +2067,7 @@ "| Tier | Purpose | SQL Equivalent |\n", "|------|---------|----------------|\n", "| `Lookup` | Reference data, parameters | Regular table |\n", - "| `Manual` | Data inserted directly from outside the pipeline (people or ingestion scripts) | Regular table |\n", + "| `Manual` | Data inserted directly from outside the pipeline (people or entry scripts) | Regular table |\n", "| `Imported` | Data from external sources (files, instruments, databases), populated by `make()` | Regular table + trigger |\n", "| `Computed` | Derived results | Materialized view + trigger |\n", "\n", diff --git a/src/tutorials/basics/01-first-pipeline.ipynb b/src/tutorials/basics/01-first-pipeline.ipynb index 031d3f31..c5d83802 100644 --- a/src/tutorials/basics/01-first-pipeline.ipynb +++ b/src/tutorials/basics/01-first-pipeline.ipynb @@ -13,7 +13,7 @@ "- Use the four core operations: restriction, projection, join, aggregation\n", "- Understand the schema diagram\n", "\n", - "We'll work with **Manual tables** only—tables where you enter data directly. Later tutorials introduce automated computation.\n", + "We'll work with **Manual tables** only—tables whose rows are written from outside the table, here by you. A Manual table is not necessarily hand-filled; an instrument or an entry script writes to one the same way. Later tutorials introduce automated computation.\n", "\n", "> **Database Backend:** All tutorials work identically on **MySQL** and **PostgreSQL** (PostgreSQL support added in DataJoint Python 2.1). The examples shown here were executed on PostgreSQL, but you can follow along using either backend.\n", "\n", diff --git a/src/tutorials/basics/02-schema-design.ipynb b/src/tutorials/basics/02-schema-design.ipynb index b93b334a..2f8fc670 100644 --- a/src/tutorials/basics/02-schema-design.ipynb +++ b/src/tutorials/basics/02-schema-design.ipynb @@ -48,20 +48,22 @@ "source": [ "## Table Tiers\n", "\n", - "DataJoint has four table tiers, each serving a different purpose:\n", + "DataJoint has four table tiers, classified by what puts rows in them.\n", "\n", - "| Tier | Class | Purpose | Data Entry |\n", - "|------|-------|---------|------------|\n", - "| **Manual** | `dj.Manual` | Core experimental data | Inserted directly by operators, instruments, or ingestion scripts |\n", - "| **Lookup** | `dj.Lookup` | Reference/configuration data | Pre-populated, rarely changes |\n", - "| **Imported** | `dj.Imported` | Data from external sources (files, instruments, databases) | Auto-populated via `make()` |\n", - "| **Computed** | `dj.Computed` | Derived/processed data | Auto-populated via `make()` |\n", + "| Tier | Class | What puts rows in it | Typical use |\n", + "|------|-------|----------------------|-------------|\n", + "| **Manual** | `dj.Manual` | A writer outside the table — an operator, an instrument, an entry script | Core experimental data |\n", + "| **Lookup** | `dj.Lookup` | The schema definition itself — the table's committed `contents` | Reference and configuration data |\n", + "| **Imported** | `dj.Imported` | The table's own `make()`, reading an external source | Files, instruments, external databases |\n", + "| **Computed** | `dj.Computed` | The table's own `make()`, deriving from other tables | Derived and processed data |\n", "\n", - "**Manual** tables are not necessarily populated by hand—they contain data entered into the pipeline by operators, instruments, or ingestion scripts using `insert` commands. In contrast, **Imported** and **Computed** tables are auto-populated by calling the `.populate()` method, which invokes the `make()` callback for each missing entry.\n", + "Rows reach a **Manual** table from outside it, through `insert` — a person typing, an instrument, or a nightly entry script all write the same way. **Imported** and **Computed** are the tiers the framework fills: calling `.populate()` invokes `make()` for each missing entry.\n", + "\n", + "See [the Relational Workflow Model](../../../explanation/relational-workflow-model/#workflow-steps-and-table-tiers) for the full axis, including why `Part` is not a fifth tier.\n", "\n", "### Manual Tables\n", "\n", - "Manual tables store data that is inserted directly—the starting point of your pipeline." + "Manual tables store rows written from outside the table—the starting point of your pipeline." ] }, { @@ -1420,7 +1422,7 @@ "- Keep keys minimal but sufficient for uniqueness\n", "\n", "### 2. Use Appropriate Table Tiers\n", - "- **Manual**: Data inserted directly — by operators, instruments, or ingestion scripts\n", + "- **Manual**: Data inserted directly — by operators, instruments, or entry scripts\n", "- **Lookup**: Configuration, parameters, reference data\n", "- **Imported**: Data read from external sources (recordings, images, instruments)\n", "- **Computed**: Derived analyses and summaries\n", diff --git a/src/tutorials/domain/calcium-imaging/calcium-imaging.ipynb b/src/tutorials/domain/calcium-imaging/calcium-imaging.ipynb index 288648cd..ea642dcc 100644 --- a/src/tutorials/domain/calcium-imaging/calcium-imaging.ipynb +++ b/src/tutorials/domain/calcium-imaging/calcium-imaging.ipynb @@ -1517,7 +1517,7 @@ "metadata": {}, "source": [ "**Legend:**\n", - "- **Green rounded boxes**: Manual tables (data entered at runtime — by a person, an instrument, or an ingestion script)\n", + "- **Green rounded boxes**: Manual tables (data entered at runtime — by a person, an instrument, or an entry script)\n", "- **Gray rounded boxes**: Lookup tables (parameters, part of the schema definition)\n", "- **Blue ellipses**: Imported tables (data from files)\n", "- **Orange ellipses**: Computed tables (derived from other tables)\n", diff --git a/src/tutorials/examples/blob-detection.ipynb b/src/tutorials/examples/blob-detection.ipynb index e0c6d643..8cbfd375 100644 --- a/src/tutorials/examples/blob-detection.ipynb +++ b/src/tutorials/examples/blob-detection.ipynb @@ -637,7 +637,7 @@ "metadata": {}, "source": [ "The diagram shows:\n", - "- **Green** = Manual tables (data inserted directly, manually or via ingestion)\n", + "- **Green** = Manual tables (data inserted directly, manually or by an entry script)\n", "- **Gray** = Lookup tables (reference data)\n", "- **Red** = Computed tables (derived data)\n", "- **Edges** = Dependencies (foreign keys), always flow top-to-bottom"