Build multi-step AI workflows with schema-guided reasoning. Works with Ollama, LMStudio, OpenAI, OpenRouter, Gemini, and all the latest models for structured generation, chaining, and data processing.
- Ollama - Local model inference
- LM Studio - Local model management
- OpenAI - Cloud-based models
- OpenRouter - Multi-provider access
- Gemini - Google DeepMind's multimodal LLMs
- JSON Schema Validation - Structured output with type safety (YAML-native or JSON string formats)
- Text Generation - Flexible content creation
- Explicit Iteration -
count: Nfor generators,forEach: stepto run once per row of an earlier step; reference the current row as{{.item.field}} - Parallel Rows -
concurrency: Ngenerates rows of a prompt step in parallel while keeping output in row order - Native Template Values - referenced values keep their JSON types:
{{range .item.companies}},{{len .item.tags}},{{if .item.isActive}}all work; arrays still print asa, band numbers verbatim - Schema-Guided Reasoning (SGR) - Guide LLMs through systematic analysis using structured schemas
- Image Analysis - Visual model integration
- CLI Integration - Use any command-line tool as a step
- Dataset Loading - Import from Huggingface
- Transform Steps - Embedded jq (via gojq): filter, reshape, and fan out data between steps — no external binary needed
- Environment Variables - Dynamic configuration with
$VARsyntax - Retry Logic - Smart error handling and recovery
brew tap mirpo/homebrew-tools
brew install datamaticgo install github.com/mirpo/datamatic@latestgit clone https://github.com/mirpo/datamatic.git
cd datamatic
make build- Synthetic Data Generation - Create training datasets for fine-tuning LLMs
- Document Classification - Systematic analysis with structured reasoning
- SQL Query Generation - Chain-of-thought reasoning for complex queries
- Multi-step Processing Pipelines - CV analysis, data transformation, content generation
- Vision Workflows - Image analysis combined with text generation
- Data Integration - Combine HuggingFace datasets with LLM processing
Create a configuration file and run datamatic:
# config.yaml
version: 1.0
steps:
- name: generate_titles
model: ollama:llama3.2
count: 5 # generate 5 rows
prompt: Generate a catchy news title
jsonSchema:
type: object
properties:
title:
type: string
tags:
type: array
items:
type: string
required:
- title
- tags
additionalProperties: false
- name: analyze_title
model: ollama:llama3.2
forEach: generate_titles # one iteration per generated title
prompt: |
Analyze this news title and provide sentiment and category analysis:
Title: {{.item.title}}
jsonSchema: |
{
"type": "object",
"properties": {
"sentiment": {"type": "string", "enum": ["positive", "negative", "neutral"]},
"category": {"type": "string", "description": "News category"},
"clickbait_score": {"type": "number", "minimum": 0, "maximum": 10}
},
"required": ["sentiment", "category", "clickbait_score"]
}# Generate data
datamatic --config config.yaml
# With debug output
datamatic --config config.yaml --verbose --log-pretty
# Check a config without running anything (great as a CI step for
# committed workflows): parses, preprocesses and validates — schemas,
# cross-step references, jq programs — and exits non-zero on any error
datamatic validate --config config.yamlOther providers:
- OpenAI:
model: openai:gpt-4o-mini+export OPENAI_API_KEY=sk-... - OpenRouter:
model: openrouter:meta-llama/llama-3.2-3b+export OPENROUTER_API_KEY=sk-... - Gemini:
model: gemini:gemini-2.0-flash+export GEMINI_API_KEY=...
Rows of a prompt step are independent, so they can be generated in parallel:
steps:
- name: analyze
model: openai:gpt-4o-mini
forEach: documents
concurrency: 5 # up to 5 rows generated at once (default: 1)- Applies to prompt steps only (
countorforEach); using it on transform or shell steps is a config error. - Output stays in row order regardless of which request finishes first, so datasets remain deterministic.
- Raise it for cloud providers, which handle many parallel requests. Keep it low (or
1) for a single local GPU — Ollama/LM Studio serve only a few requests at a time, so a high value won't help and may thrash.
Reshape, filter, and fan out data between steps with embedded jq (via gojq — no external binary needed):
steps:
- name: picked
from: source_step
jq: 'select(.score > 5) | {q: .question, a: .answer}'
limit: 100from— source step; the jq program sees each row's value (for prompt steps: theresponse)jq— any jq program; emitting multiple values fans out (1 row → N rows),select()filters rows outcollect: true— fan-in: the program runs once over an array of all source rows (unique,group_by,sort_byacross the whole dataset)sourceFormat: json— the source file is a single JSON value (e.g. a pretty-printed array from an API dump) instead of JSONL$parent— per-row programs can reach the source row's lineage as$parent.step.field(e.g. carry the original chunk while fanning out extracted questions); not available withcollect, where there is no single parent rowlimit— optional cap on output rows
Always wrap jq programs in single quotes: unquoted YAML silently truncates at #, misparses {...} object construction, and jq's own strings use double quotes anyway.
jq programs are validated when the config loads. Transform steps run instantly, produce regular JSONL, and don't trigger the external-CLI warning. See the dataset-pipeline example, which uses fan-out and fan-in.
Process your own data end to end — no shell glue:
steps:
- name: leads
read: leads.csv # local files → rows
- name: classified
forEach: leads
prompt: Classify {{.item.company}}
jsonSchema: { ... }
- name: report
from: classified
write: enriched.csv # rows → a deliverable fileread:turns local files into rows. Format is inferred from the path (overridable withformat:):- a glob / directory /
.txt/.md→ one row per file:{path, name, content} .csv/.tsv→ one row per record (columns become fields).jsonl→ one row per line
- a glob / directory /
write:exports a step's rows to a file, format inferred from the extension:.csv,.json(array),.md(table), or.jsonl. It's terminal and doesn't change the intermediate JSONL that other steps read.image:on a prompt step attaches a file as a vision image, e.g.image: "{{.item.path}}"afterread-ing a folder of images.
A write step's source picks its mode. from: writes one file with every row — a dataset or report. forEach: writes one file per row — a folder of documents:
- name: reply_digest
from: drafts # → one replies.md table of all drafts
write: replies.md
- name: reply_files
forEach: drafts # → one file per draft
write: replies/{{.item.subject}}.md # path is a template, rendered per row
content: "{{.item.reply}}" # body is raw text, not JSON- The path template may nest (
write: out/{{.item.kind}}/{{.item.id}}.md); missing directories are created. content:is the file body as raw text — that is what makes the output a document rather than a record. Omit it and the row is serialized by extension instead (.json,.csv, …), one row per file.- Values interpolated into the path are made filename-safe, so a
/or:inside your data can't redirect the file — only slashes you write in the template create directories. A name that renders empty falls back to the row number, and two rows producing the same name get numbered (-2) instead of overwriting each other.
Where paths point. Inputs travel with the workflow; everything generated lands in the output folder:
| Path | Relative to | Absolute |
|---|---|---|
read: (input) |
the config file's directory — so a workflow runs from any working directory | used as-is |
write: (deliverable) |
the output folder | used as-is — this is how you publish outside it |
write: per-row template |
the output folder, after rendering | used as-is |
| intermediate JSONL | the output folder | — |
--output (flag) |
the working directory | used as-is |
The output folder is reused and overwritten on each run — no dataset_v1, dataset_v2 copies, so deliverable paths stay stable. For run history, point --output somewhere per-run (e.g. --output runs/$(date +%F-%H%M)). dataset* is gitignored.
One exception: a shell step with an explicit workDir writes its outputFilename into that directory (relative to the output folder, or wherever an absolute workDir points), since the command needs to produce the file in its own working directory.
See the csv-enrichment and process-my-files examples.
Shape the JSON schema to steer how the model reasons, not just what it returns. Three patterns (background):
- Cascade — put reasoning before the conclusion (a
reasoningfield, or asteps[]array, ahead of the answer). The workhorse; works well even on small local models. See sgr-reasoning, document-classification, inbox-triage. - Routing — a discriminated union (
anyOfof object branches, each with aconstdiscriminator) makes the model pick one branch and fill only its fields; datamatic validates the union and every branch for strict output. Branch-choice accuracy needs a capable model — small local models (≤3B) reliably mis-route, so use a cloud model for real routing. - Cycle — an
arrayof a repeated sub-schema (optionally bounded withminItems/maxItems) emits N reasoning items. Bounds are honored by Ollama's grammar but rejected by OpenAI strict mode.
Prompt steps send the schema as strict structured output (all properties required, additionalProperties: false); datamatic validate flags schemas that break those rules before you hit the API.
Configure your pipelines dynamically using $VAR syntax:
version: 1.0
envVars:
- PROVIDER
- MODEL
steps:
- name: generate
model: $PROVIDER:$MODEL
prompt: Generate a creative storyPROVIDER=ollama MODEL=llama3.2 datamatic --config config.yamlVariables listed in envVars are validated before execution (fail-fast). See env-and-workdir example for more details.
Datamatic outputs structured data in JSONl format:
type LineEntity struct {
ID string `json:"id"`
Format string `json:"format"`
Prompt string `json:"prompt"`
Response interface{} `json:"response"`
Values map[string]promptbuilder.ValueShort `json:"values,omitempty"`
}- Format:
textorjson - Response: Generated content (text string or JSON object)
- Values: Linked step values for traceability
Text line:
{
"id":"38082542-f352-44d2-88e9-6d68d28dcac4"
"format":"text",
"prompt":"Generate a catchy and one unique news title. Come up with a wildly different and surprising news headline. Return only one news title per request, without any extra thinking.",
"response":"BREAKING: Giant Squid Found Wearing Tiny Top Hat and monocle in Remote Arctic Location"
}JSON line:
{
"id":"cc437b10-63c6-443a-9b3e-a7d6c51fc0a0",
"format":"json",
"prompt":"Provide up-to-date information about a randomly selected country, including its name, population, land area, UN membership status, capital city, GDP per capita, official languages, and year of independence. Return the data in a structured JSON format according to the schema below.",
"response":{"capitalCity":"Bishkek","gdpPerCapita":1700,"independenceYear":1991,"isUNMember":true,"languages":["Kyr Kyrgyz","Russian"],"name":"Kyrgyzstan","population":6184000,"totalCountryArea":199912}
}With values from linked steps:
{
"id":"dc140355-6c41-4ce7-9127-b8145cf1a23e",
"format":"text",
"prompt":"Write nice tourist brochure about country Kyrgyzstan (a UN member state), which capital is Bishkek, area 199912, independenceYear: 1991 and official languages (2 total): Kyrgyz, Russian.",
"response":"**Discover the Hidden Gem of Central Asia: Kyrgyzstan**\n\nTucked away in the heart of Central Asia, Kyrgyzstan is a land of breathtaking beauty, rich history, and warm hospitality. Our capital city, Bishkek, is a bustling metropolis surrounded by the stunning Tian Shan mountains, waiting to be explored.\n\n**A Brief History**\n\nKyrgyzstan gained its independence on August 31, 1991...",
"values":{".about_country.capitalCity":{"id":"cc437b10-63c6-443a-9b3e-a7d6c51fc0a0","value":"Bishkek"},".about_country.independenceYear":{"id":"cc437b10-63c6-443a-9b3e-a7d6c51fc0a0","value":1991},".about_country.isUNMember":{"id":"cc437b10-63c6-443a-9b3e-a7d6c51fc0a0","value":true},".about_country.languages":{"id":"cc437b10-63c6-443a-9b3e-a7d6c51fc0a0","value":["Kyrgyz","Russian"]},".about_country.name":{"id":"cc437b10-63c6-443a-9b3e-a7d6c51fc0a0","value":"Kyrgyzstan"},".about_country.totalCountryArea":{"id":"cc437b10-63c6-443a-9b3e-a7d6c51fc0a0","value":199912}}
}datamatic [OPTIONS] # run the workflow
datamatic validate [OPTIONS] # check the config and exit (0 = valid)
Options:
-config string
Config file path
-http-timeout int
HTTP timeout: 0 - no timeout, if number - recommended to put high on poor hardware (default 300)
-log-pretty
Enable pretty logging, JSON when false (default true)
-output string
Output folder path, relative to the working directory
(default: 'dataset' next to the config file)
-validate-response
Validate JSON response from server to match the schema (default true)
-verbose
Enable DEBUG logging level
-version
Get current version of datamaticChoosing the output folder — first match wins:
--output <path>— relative to the working directory, absolute used as-isoutput:in the config — relative to the config file's directory, so it travels with the workflow- otherwise
datasetnext to the config file — the same workflow lands in the same place no matter where it was launched from
datamatic --config flows/triage/config.yaml # → flows/triage/dataset/
datamatic --config flows/triage/config.yaml --output ./today # → ./today/
datamatic --config flows/triage/config.yaml --output /srv/runs/a # → /srv/runs/a/See examples/v1/ for the full feature matrix. Start with basics, then linked-steps.
| Example | Features shown | Backend |
|---|---|---|
| basics | text generation, JSON schema | Ollama |
| linked-steps | step chaining, native template values (if/range/len) |
Ollama |
| structured-extraction | nested schema, both schema formats (YAML/JSON-string), native templates | Ollama |
| Example | Features shown | Backend |
|---|---|---|
| dataset-pipeline | transform fan-out, fan-in (collect, $parent), rating pipeline |
Ollama |
| sgr-reasoning | schema-guided reasoning, sourceFormat: json |
Ollama |
| document-classification | SGR classification, collect fan-in QA |
Ollama |
| document-qa | RAG-style Q&A from a document, rating filter | Ollama |
| Example | Features shown | Backend |
|---|---|---|
| process-my-files | read local files (glob/dir/CSV/JSONL) into rows |
Ollama |
| csv-enrichment | read CSV → LLM enrich → write CSV (full office loop) |
Ollama |
| external-data | HuggingFace download + transform + shell tools | Ollama |
| env-and-workdir | env vars, workDir, $PROVIDER, DuckDB |
Ollama |
| vision | image → structured output (imagePath) |
Ollama, LM Studio |
| Example | Features shown | Backend |
|---|---|---|
| cloud-providers | provider selection, concurrency, retryConfig |
OpenAI / OpenRouter / Gemini |