[TOC]
The In-Memory API is the fastest way to experiment with DPSynth. Built on top of Pandas and NumPy, this interface is designed for researchers, rapid prototypers, and software engineers operating on datasets that comfortably fit within a single machine's RAM.
The primary entry point for in-memory synthesis is
dpsynth.TabularConfig. It accepts a dictionary of attribute domains and
mechanism options, is calibrated with a privacy budget to produce a
dpsynth.TabularMechanism, and generates a fully synthetic, differentially
private DataFrame matching the exact schema and data types of your input.
import dpsynth
from dpsynth import discrete_mechanisms
import numpy as np
import pandas as pd
config = dpsynth.TabularConfig(
domains=domains,
discrete_mechanism=discrete_mechanisms.MSTConfig(),
)
mechanism = config.calibrate(epsilon=1.0, delta=1e-6)
result = mechanism(np.random.default_rng(), sensitive_df)
synthetic_df = result.synthetic_dataWhen initializing dpsynth.TabularConfig:
domains: Mapping of column names to domain specifications (CategoricalAttribute,NumericalAttribute, orOpenSetCategoricalAttribute). Every key must exist indata.columns.discrete_mechanism: Configuration object specifying which DP synthesis mechanism to run (e.g.,MSTConfig(),AIMConfig(),IndependentConfig()).numerical_bins: Number of equal-frequency quantile buckets used to discretize continuous numerical columns (default:32).init_budget_fraction: Fraction of total(epsilon, delta)budget allocated for per-column initialization such as bounds computation and partition selection (default:0.1).cross_attribute_constraints: Optional sequence of constraints to enforce on generated data.
When calling config.calibrate(...):
epsilon,delta: Total differential privacy budget parameters. Returns a runnableTabularMechanism.
Here is a complete, self-contained Python script demonstrating how to specify a
domain, set up a TabularConfig, calibrate the mechanism with a privacy budget,
load sensitive data, synthesize records, and print the first few rows.
import dpsynth
from dpsynth import discrete_mechanisms
from dpsynth import domain
import numpy as np
import pandas as pd
# 1. Domain Specification: Define the schema of the tabular dataset
attribute_domains = {
"age": domain.NumericalAttribute(lower_bound=18, upper_bound=90),
"workclass": domain.CategoricalAttribute(
allowed_values=["Private", "Self-emp", "Gov", "Other"]
),
"education": domain.CategoricalAttribute(
allowed_values=["HS-grad", "Bachelors", "Masters", "PhD"]
),
}
# 2. Setup Config: Configure synthesizer with domain and mechanism choices
config = dpsynth.TabularConfig(
domains=attribute_domains,
discrete_mechanism=discrete_mechanisms.MSTConfig(),
numerical_bins=16,
)
# 3. Calibrate Mechanism: Allocate privacy budget to get runnable mechanism
mechanism = config.calibrate(epsilon=1.0, delta=1e-5)
# 4. Load Data: Create sensitive input DataFrame matching the domain schema
sensitive_df = pd.DataFrame({
"age": [25, 42, 30, 55, 62, 29, 38, 47, 51, 33],
"workclass": [
"Private",
"Gov",
"Private",
"Self-emp",
"Other",
"Private",
"Gov",
"Private",
"Self-emp",
"Private",
],
"education": [
"Bachelors",
"Masters",
"HS-grad",
"PhD",
"HS-grad",
"Bachelors",
"HS-grad",
"Masters",
"Bachelors",
"HS-grad",
],
})
# 5. Synthesize Data: Run the calibrated mechanism on the sensitive data
rng = np.random.default_rng(seed=42)
result = mechanism(rng, sensitive_df)
synthetic_df = result.synthetic_data
# 6. Print the first few rows of the generated synthetic dataset
print("Generated Synthetic Data:")
print(synthetic_df.head())For immediate execution without writing custom Python scripts, use the
standalone
binary bin/main.py.
It provides command-line flags for all standard configuration parameters.
python3 bin/main.py \
--dataset=/path/to/dataset.csv \
--domain=/path/to/domain.yaml \
--epsilon=1.0 \
--delta=1e-8 \
--mechanism=mst \
--seed=12345 \
--output_path=/tmp/synthetic_output.csv--dataset: Path to the input CSV file. (Supports standard CSV parsing arguments via--read_csv_args).--domain: Path to the YAML domain specification file.--epsilon,--delta: Total DP privacy budget.--mechanism: Supported options aremst,aim, andindependent.--seed: Integer seed for reproducible randomness across DP sampling and PGM inference.--output_path: Destination filepath where the synthetic CSV will be written.
When you configure and run TabularConfig, the library performs the following
single-machine pipeline:
- Discretization: Continuous numerical columns are bucketed into
numerical_binsquantiles usingpipeline_dp.LocalBackend. Open-set strings are evaluated via DP partition selection. - Integer Encoding: All columns are mapped to dense integer indices
[0, K-1]. - Domain Compression: DPSynth measures 1-way marginals with Gaussian noise
and merges rare categories into an
"Other"bucket, producing an un-noised discrete dataset (mbi.Dataset). - Mechanism Execution: Calls the configured discrete mechanism (
AIM,MST, etc.) on the discrete dataset. The mechanism fits a Markov Random Field (mbi.MarkovRandomField) via Private-PGM mirror descent. - Sampling & Inversion: Samples synthetic integer records from the
graphical model, unpacks
"Other"categories, and inverts the integer encoding back to original Pandas dtypes (strings, integers, floating points).