Skip to content
Open
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
35 commits
Select commit Hold shift + click to select a range
c079048
add a plan for replacing detection validity score
memadi-nv Jun 24, 2026
e7817f7
nit
memadi-nv Jun 24, 2026
e4ab7c6
address feedback
memadi-nv Jun 24, 2026
dca7462
add missing url to the plan
memadi-nv Jun 24, 2026
fcb3feb
nit
memadi-nv Jun 24, 2026
10def26
update plan- keep precision add recall
memadi-nv Jun 26, 2026
d6ded84
add entity coverage
memadi-nv Jun 29, 2026
5ed4b46
optional detection validity in rewrite
memadi-nv Jul 1, 2026
2c5948f
add changes fro nemotron super test and entity coverage prompt
memadi-nv Jul 2, 2026
277a7a2
include strict entity protection in evaluate
memadi-nv Jul 9, 2026
6618369
add post processing for entity coverage
memadi-nv Jul 14, 2026
4587547
add nemotron ultra as entity coverage judge
memadi-nv Jul 14, 2026
da98003
explore entity coverage judge prompt
memadi-nv Jul 16, 2026
b89481b
nit
memadi-nv Jul 16, 2026
fddb437
improve entity coverage candidate extraction
memadi-nv Jul 17, 2026
6f8d2e0
add data summary to entity coverage
memadi-nv Jul 17, 2026
f706a55
add evaluation logs
memadi-nv Jul 18, 2026
474d452
remove plan
memadi-nv Jul 18, 2026
2c554cb
change entity coverage to nemotron super
memadi-nv Jul 18, 2026
07ae62f
update docs
memadi-nv Jul 19, 2026
dfe3086
update docs
memadi-nv Jul 19, 2026
f999832
nit
memadi-nv Jul 19, 2026
63b3786
fix: restore check_rewrite_judge validation in evaluate() rewrite path
memadi-nv Jul 20, 2026
155b91e
address greptile feedback
memadi-nv Jul 27, 2026
c8392f5
update model name to match build
memadi-nv Jul 29, 2026
9fabaaf
Apply suggestions from code review in evaluation.md
memadi-nv Jul 29, 2026
8238f76
edit wording-not mentioning customer facing
memadi-nv Jul 29, 2026
a9d95b1
restrict stopword coverage to leading articles only
memadi-nv Jul 29, 2026
0573f4c
update prompt-less duplication
memadi-nv Jul 29, 2026
648deba
update prompt for strict entity protection
memadi-nv Jul 29, 2026
356e355
replace LLMTextColumnConfig with LLMStructuredColumnConfig
memadi-nv Jul 29, 2026
76735f7
add deterministic label filter in postprocessing step
memadi-nv Jul 29, 2026
4f788a3
nit - enable set of value and label for total entity count
memadi-nv Jul 29, 2026
f647b9f
nit
memadi-nv Jul 29, 2026
cdb4a49
nit
memadi-nv Jul 29, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
6 changes: 4 additions & 2 deletions AGENTS.md
Original file line number Diff line number Diff line change
Expand Up @@ -75,9 +75,11 @@ input_df
RepairWorkflow → _rewritten_text (failing rows only)
→ output: {text_col}_rewritten, utility_score, leakage_mass, needs_human_review, …

Later, `RewriteWorkflow.evaluate()` can run the non-critical judge steps:
Later, `Anonymizer.evaluate()` runs the non-critical judge steps:
→ EntityCoverageWorkflow → entity_coverage, leaked_entities (always)
→ FinalJudgeWorkflow → judge_evaluation (always)
→ DetectionJudgeWorkflow → detection_valid, detection_invalid_entities
→ FinalJudgeWorkflow → judge_evaluation
(only when EvaluateConfig.compute_detection_validity=True)
```

Records with no detected entities skip all LLM sub-workflows and pass through with default metrics (utility=1.0, leakage=0.0).
Expand Down
64 changes: 54 additions & 10 deletions docs/concepts/evaluation.md
Original file line number Diff line number Diff line change
Expand Up @@ -45,7 +45,41 @@ with open("/tmp/preview.pkl", "rb") as f:
evaluated = anonymizer.evaluate(loaded)
```

Four LLM judges run per record: one that scores detection quality and three that score replacement quality (Substitute mode only).
The active judges depend on the anonymization strategy and `EvaluateConfig`:

| Evaluation | Applies to |
|------------|------------|
| Entity Coverage | All anonymization strategies |
| Detection Validity | Optional; enable with EvaluateConfig(compute_detection_validity=True) |
| Type Fidelity, Attribute Fidelity, Relational Consistency** | Substitute mode only |

---

### Entity Coverage

> "Which sensitive values from the original text survived into the anonymized output?"

Entity coverage is the primary residual-leakage metric and **always runs**. The judge scans the original text independently and reports any in-scope sensitive values that were **not** removed or replaced. A deterministic postprocessing step then drops candidates that are already accounted for by the Anonymizer's final entities (exact, sub-span, or composite matches), so only genuine leaks remain.

The judge is scoped and contextualized by the same signals used during anonymization:

- **`entity_labels`** — the detection taxonomy in scope; the judge only reports values whose type falls within it.
- **`data_summary`** — used purely to interpret literal values and their semantic types, never to invent entities absent from the text.
- **`strict_entity_protection`** — (rewrite only) when enabled, the judge also reports inferable/indirect sensitive values, not just literal identifiers.

| Output column | Type | Description |
|---|---|---|
| `entity_coverage` | `float \| None` | `n_final / (n_final + n_leaked)` — fraction of sensitive values that were removed. `1.0` means nothing leaked; `None` if the judge was unavailable. |
| `leaked_entities` | `list` | Each surviving value with its `value`, `label`, and one-sentence `reasoning`. Empty when nothing leaked. |

**Special values:**

| Scenario | `entity_coverage` | `leaked_entities` |
|---|---|---|
| No sensitive values in the original text | `1.0` | `[]` |
| Judge ran and found no surviving values | `1.0` | `[]` |
| Judge ran and found surviving values | 0–1 fraction | populated |
| Judge call failed or returned a malformed response | `None` | `[]` |

---

Expand All @@ -55,7 +89,7 @@ Four LLM judges run per record: one that scores detection quality and three that

> "Are the detected entities actually correct (value, label) pairs in context?"

This judge runs regardless of which replace mode was used. It looks at each detected span and flags:
Detection validity is disabled by default. To enable it, set `compute_detection_validity=True` in `EvaluateConfig`. This metric is intended for internal model and threshold evaluation. It examines each detected span and flags:

- **false_positive** — the span is not actually identifying or sensitive in this context (common word, generic phrase, boilerplate).
- **wrong_label** — the span is sensitive but the label sits in a clearly different domain (e.g. a company name labeled `first_name`). Sibling labels within the same broad domain are treated as valid.
Expand Down Expand Up @@ -138,10 +172,11 @@ For a tabular overview across all records:
```python
evaluated.dataframe[
[
"detection_valid",
"entity_coverage",
"type_fidelity_valid",
"attribute_fidelity_valid",
"relational_consistency_valid",
# "detection_valid" — present only if compute_detection_validity=True
]
]
```
Expand All @@ -152,14 +187,15 @@ Use `trace_dataframe` for the full internal trace including raw judge outputs.

### Model roles

All four judges default to `gpt-oss-120b`. Defaults are defined in [`evaluate.yaml`](https://github.com/NVIDIA-NeMo/Anonymizer/blob/main/src/anonymizer/config/default_model_configs/evaluate.yaml). Override them by passing a `model_configs` YAML to `Anonymizer(model_configs=...)` — see [Models](models.md) for the full override pattern.
The entity coverage judge defaults to `nemotron-super`; the other replace-evaluation judges default to `gpt-oss-120b`. Defaults are defined in [`evaluate.yaml`](https://github.com/NVIDIA-NeMo/Anonymizer/blob/main/src/anonymizer/config/default_model_configs/evaluate.yaml). Override them by passing a `model_configs` YAML to `Anonymizer(model_configs=...)` — see [Models](models.md) for the full override pattern.

The four roles are `detection_validity_judge`, `replace_type_fidelity_judge`, `replace_attribute_fidelity_judge`, and `replace_relational_consistency_judge`.
The roles are `entity_coverage_judge`, `detection_validity_judge`, `replace_type_fidelity_judge`, `replace_attribute_fidelity_judge`, and `replace_relational_consistency_judge`.

```yaml
# my_models.yaml
selected_models:
evaluate:
entity_coverage_judge: your-model-alias
detection_validity_judge: your-model-alias
replace_type_fidelity_judge: your-model-alias
replace_attribute_fidelity_judge: your-model-alias
Expand All @@ -174,7 +210,7 @@ Rewrite evaluation has two layers:

1. **Automatic (always runs)** — leakage mass, utility score, weighted leakage rate, and `needs_human_review` are computed as part of every `run()` / `preview()` call. See [Rewrite](rewrite.md) for the repair loop and output columns.

2. **Post-hoc LLM judges (optional)** — call `Anonymizer.evaluate()` on a completed rewrite result to add the entity detection judge and three holistic quality rubrics.
2. **Post-hoc LLM judges (optional)** — call `Anonymizer.evaluate()` on a completed rewrite result to add the entity coverage judge (always), three holistic quality rubrics (always), and the detection validity judge (only with `EvaluateConfig(compute_detection_validity=True)`).

```python
from anonymizer import Anonymizer, AnonymizerConfig, AnonymizerInput, Rewrite
Expand Down Expand Up @@ -207,9 +243,15 @@ evaluated = anonymizer.evaluate(loaded)

---

### Entity Coverage

Same judge as in replace mode — see [Entity Coverage](#entity-coverage) above. It **always runs** and emits `entity_coverage` and `leaked_entities`. In rewrite mode the judge additionally honours `strict_entity_protection`: when the rewrite config enables it, inferable/indirect sensitive values are reported as leaks, not just literal identifiers.

---

### Entity Detection Judge

Same judge as in replace mode — see [Entity Detection Judge](#entity-detection-judge) above. In rewrite mode, `detection_valid` is returned as a **0–1 fraction** (the share of detected entities that passed), rather than a boolean. A value of `1.0` means all detections are valid; lower values mean more entities were flagged — the value itself is the fraction that passed.
Same judge as in replace mode — see [Entity Detection Judge](#entity-detection-judge) above, and like there it is **opt-in** via `EvaluateConfig(compute_detection_validity=True)` (off by default). In rewrite mode, `detection_valid` is returned as a **0–1 fraction** (the share of detected entities that passed), rather than a boolean. A value of `1.0` means all detections are valid; lower values mean more entities were flagged — the value itself is the fraction that passed.

| Output column | Type | Description |
|---|---|---|
Expand Down Expand Up @@ -282,7 +324,7 @@ The three rubric scores are stored together under the `judge_evaluation` column

### Reading rewrite evaluation results

`display_record()` renders a formatted per-record view that includes the detection validity fraction and all three judge rubrics alongside the rewritten text:
`display_record()` renders a formatted per-record view that includes entity coverage, all three judge rubrics, and (when enabled) the detection validity fraction alongside the rewritten text:

```python
evaluated.display_record(0)
Expand All @@ -291,7 +333,8 @@ evaluated.display_record(0)
For a tabular overview across all records:

```python
evaluated.dataframe[["detection_valid", "judge_evaluation"]]
evaluated.dataframe[["entity_coverage", "judge_evaluation"]]
# add "detection_valid" if you evaluated with compute_detection_validity=True
```

Use `trace_dataframe` for the full internal trace including raw judge outputs.
Expand All @@ -300,12 +343,13 @@ Use `trace_dataframe` for the full internal trace including raw judge outputs.

### Model roles

The rewrite quality judge defaults to `nemotron-30b-thinking`. The detection validity judge shares the `detection_validity_judge` role used by replace evaluation. Defaults are defined in [`evaluate.yaml`](https://github.com/NVIDIA-NeMo/Anonymizer/blob/main/src/anonymizer/config/default_model_configs/evaluate.yaml). Override them via `model_configs`:
The rewrite quality judge defaults to `nemotron-30b-thinking` and the entity coverage judge to `nemotron-super`. The detection validity judge shares the `detection_validity_judge` role used by replace evaluation. Defaults are defined in [`evaluate.yaml`](https://github.com/NVIDIA-NeMo/Anonymizer/blob/main/src/anonymizer/config/default_model_configs/evaluate.yaml). Override them via `model_configs`:

```yaml
# my_models.yaml
selected_models:
evaluate:
entity_coverage_judge: your-model-alias
detection_validity_judge: your-model-alias
rewrite_judge: your-model-alias
```
1 change: 1 addition & 0 deletions docs/concepts/models.md
Original file line number Diff line number Diff line change
Expand Up @@ -26,6 +26,7 @@ export NVIDIA_API_KEY="your-nvidia-api-key"
| `gliner-pii-detector` | [`nvidia/gliner-pii`](https://build.nvidia.com/nvidia/gliner-pii) | Entity detection (NER) |
| `gpt-oss-120b` | [`openai/gpt-oss-120b`](https://build.nvidia.com/openai/gpt-oss-120b) | Detection validation & augmentation, replacement, replace evaluation, rewriting |
| `nemotron-30b-thinking` | [`nvidia/nemotron-3-nano-30b-a3b`](https://build.nvidia.com/nvidia/nemotron-3-nano-30b-a3b) | Latent detection, rewrite evaluation, final judge |
| `nemotron-super` | [`nvidia/nemotron-3-super-v3`](https://build.nvidia.com/nvidia/nemotron-3-super-v3) | Entity coverage evaluation |

Each pipeline stage has a **role** mapped to one of these aliases. See the full role list in the default configs: [`detection.yaml`](https://github.com/NVIDIA-NeMo/Anonymizer/blob/main/src/anonymizer/config/default_model_configs/detection.yaml), [`replace.yaml`](https://github.com/NVIDIA-NeMo/Anonymizer/blob/main/src/anonymizer/config/default_model_configs/replace.yaml), [`rewrite.yaml`](https://github.com/NVIDIA-NeMo/Anonymizer/blob/main/src/anonymizer/config/default_model_configs/rewrite.yaml).

Expand Down
2 changes: 1 addition & 1 deletion docs/concepts/replace.md
Original file line number Diff line number Diff line change
Expand Up @@ -122,4 +122,4 @@ AnonymizerConfig(replace=Hash(algorithm="sha1", digest_length=8))

## Evaluating replace output

After running `replace`, you can score the quality of substitutions using LLM-as-judge evaluation. See [Evaluation](evaluation.md) for details on all four judges (detection validity, type fidelity, attribute fidelity, relational consistency) and how to call `Anonymizer.evaluate()`.
After running `replace`, you can score residual leakage and the quality of substitutions using LLM-as-judge evaluation. See [Evaluation](evaluation.md) for details on the always-on entity coverage judge, the three Substitute quality judges (type fidelity, attribute fidelity, relational consistency), the opt-in detection validity judge, and how to call `Anonymizer.evaluate()`.
8 changes: 5 additions & 3 deletions docs/concepts/rewrite.md
Original file line number Diff line number Diff line change
Expand Up @@ -135,9 +135,11 @@ config = AnonymizerConfig(
| `weighted_leakage_rate` | Always | Normalized leakage (0.0--1.0) relative to the maximum possible leakage mass. |
| `any_high_leaked` | Always | Whether any high-sensitivity entity leaked through. |
| `needs_human_review` | Always | Flag for records that may need manual review. |
| `detection_valid` | After `evaluate()` | Fraction of detected entities that passed the detection judge (0.0--1.0); `None` if judge unavailable. |
| `detection_invalid_entities` | After `evaluate()` | Flagged detections with value, label, and one-sentence reasoning. |
| `entity_coverage` | After `evaluate()` | Fraction of sensitive values removed (0.0--1.0); `1.0` if nothing leaked, `None` if judge unavailable. |
| `leaked_entities` | After `evaluate()` | Sensitive values that survived into the rewritten text, each with value, label, and reasoning. |
| `judge_evaluation` | After `evaluate()` | Dict with `privacy`, `quality`, and `style` rubric scores and reasoning. |
| `detection_valid` | After `evaluate(config=EvaluateConfig(compute_detection_validity=True))` | Fraction of detected entities that passed the detection judge (0.0--1.0); `None` if judge unavailable. Opt-in, off by default. |
| `detection_invalid_entities` | After `evaluate(config=EvaluateConfig(compute_detection_validity=True))` | Flagged detections with value, label, and one-sentence reasoning. |

Use `preview.trace_dataframe` for the full pipeline trace (domain, disposition, QA pairs, repair iterations).

Expand Down Expand Up @@ -187,4 +189,4 @@ Rewrite uses multiple LLM roles. All default to models in the [default config](m

## Evaluating rewrite output

After running rewrite, you can score detection quality and the holistic rewrite quality using LLM-as-judge evaluation. See [Evaluation](evaluation.md) for details on the detection judge and the three rewrite quality rubrics (privacy, quality, style), and how to call `Anonymizer.evaluate()`.
After running rewrite, you can score residual leakage and the holistic rewrite quality using LLM-as-judge evaluation. See [Evaluation](evaluation.md) for details on the always-on entity coverage judge, the three rewrite quality rubrics (privacy, quality, style), the opt-in detection judge, and how to call `Anonymizer.evaluate()`.
Loading
Loading