bmbouter Upgrade Pulp Dependencies: Design Spec
Problem
The current upgrade-deps agent has a 300-line AGENT.md that instructs an LLM
to follow an 8-phase procedure. It is non-deterministic because:
- LLMs are stochastic procedure executors. Every run produces different
shell commands, different judgment calls, different outcomes.
- The hardest task -- patch analysis -- is pure subjective judgment. The
agent must read upstream diffs, decide if a patch is "upstreamed," and
manually regenerate failing patches. This is where the agent goes off the
rails.
- No verification gates. The agent can silently skip steps, misclassify
patches, or produce broken output with no mechanism to detect it.
- No progress tracking. If the agent fails at phase 6, all work on
phases 1-5 is lost.
Design Principles
- Scripts for procedures. Exit codes for decisions. LLM only for code understanding.
- Goal-oriented loops with verification, not linear step-following.
- Always commit, even on partial failure. Partial progress > no output.
- Structured JSON artifacts between steps, not conversational memory.
Architecture
One agent, seven scripts, goal-loop orchestration.
Agent (AGENT.md ~80 lines)
|
| orchestrates via goal-loop
|
+-- 01-resolve-versions.sh -> upgrade-plan.json
+-- 02-update-requirements.sh -> requirements-diff.txt
+-- 03-analyze-patches.sh -> patch-report.json + .rej files
| |
| +-- LLM fixes conflict patches (reads .rej files)
|
+-- 04-install-and-apply.sh -> apply-results.json
+-- 05-migrate-and-test.sh -> migration-log.txt + test-results.txt
| |
| +-- LLM diagnoses failures, fixes code/patches
|
+-- 06-sync-dockerfile.sh -> dockerfile-diff.txt
+-- 07-commit-and-push.sh -> commit-sha.txt
The LLM has exactly 3 jobs:
- Fix patches with
status=conflict (read .rej files, understand intent,
rewrite hunks)
- Sanity-check
status=upstreamed verdicts (quick confirmation read)
- Diagnose test/migration failures
Everything else is a script: PyPI queries, version comparison, git tag
checkout, patch testing, patch -R --dry-run upstreaming detection, pip
install/uninstall, Dockerfile syncing, commit message generation.
Key Innovation: patch -R --dry-run for Upstreaming Detection
The current agent's #1 failure mode is the "is this patch upstreamed?"
decision. The AGENT.md gives 15 lines of subjective instructions ("check each
change individually," "verify those exact lines exist"). Different runs reach
different conclusions about the same patch.
The replacement is mechanical:
git checkout {new_tag}
patch -R --dry-run -p1 < {patch_file}
If the reverse of the patch applies cleanly (exit 0), the new version already
contains the patch's changes. No LLM judgment needed, just an exit code.
When a patch genuinely conflicts, patch --force produces .rej files that
show exactly which hunks failed and what the expected vs actual context was.
This gives the LLM a constrained problem ("fix this specific hunk against this
new context") instead of the current open-ended "read the upstream diff and
recreate the patch."
Repo Strategy
The agent uses repo_group: pulp-stack, which tells alcove to pre-clone all
upstream repos into /workspace/ before the agent starts:
/workspace/pulp-service (the main repo)
/workspace/pulpcore
/workspace/pulp_python
/workspace/pulp_container
/workspace/pulp_rpm
/workspace/pulp_maven
/workspace/pulp_gem
/workspace/pulp_npm
/workspace/pulp_hugging_face
Scripts reference these pre-cloned repos directly. No git cloning during the
run, no network failures during patch analysis.
Edge case: oras and django-storages are not in pulp-stack. Patches
targeting those packages (0010 for oras, 0031 for django-storages) are handled
by the script cloning on demand only if those packages are being upgraded.
Script Interfaces
01-resolve-versions.sh
INPUTS:
$REQUIREMENTS_FILE Path to requirements.txt
$PACKAGES (optional) JSON: {"pulpcore": "3.109.0", ...}
$WORK_DIR Working directory for artifacts
OUTPUTS:
$WORK_DIR/upgrade-plan.json
EXIT CODES:
0 Upgrades found, plan written
2 No upgrades available (clean exit)
1 Error (network, parse failure)
Output format:
{
"upgrades": [
{"package": "pulpcore", "current": "3.108.0", "target": "3.109.0",
"workspace_path": "/workspace/pulpcore"}
],
"unchanged": [
{"package": "pulp-rpm", "current": "3.35.2", "reason": "already latest"}
]
}
Uses PyPI .info.version (never release list sorting) and
packaging.version.Version for comparison. The package-to-workspace mapping
comes from lib/pkg-workspace-map.sh.
02-update-requirements.sh
INPUTS:
$WORK_DIR/upgrade-plan.json
$REQUIREMENTS_FILE
OUTPUTS:
Modified requirements.txt (in place)
$WORK_DIR/requirements-diff.txt
EXIT CODES:
0 Success
1 Error
Pure sed replacements. Verifies each target version appears in the updated
file.
03-analyze-patches.sh (the workhorse)
INPUTS:
$WORK_DIR/upgrade-plan.json
$PATCHES_DIR Path to images/assets/patches/
$DOCKERFILE Path to Dockerfile
OUTPUTS:
$WORK_DIR/patch-report.json
$WORK_DIR/rejections/{patch_name}/ (.rej files for conflicts)
EXIT CODES:
0 Analysis complete (conflicts are in the report, not errors)
1 Fatal error
Algorithm (fully deterministic):
for each patch referenced in Dockerfile (in Dockerfile order):
1. Parse diff headers -> identify target package
2. Look up package in upgrade-plan.json
3. If not upgraded -> status="skip"
4. cd /workspace/{repo}
5. TEST: forward apply at old tag
git checkout {old_tag}
patch --dry-run -p1 < {patch_file}
If fails -> status="error" (pre-existing problem)
6. TEST: forward apply at new tag
git checkout {new_tag}
patch --dry-run -p1 < {patch_file}
If succeeds -> status="clean"
7. TEST: reverse apply at new tag (upstreaming check)
patch -R --dry-run -p1 < {patch_file}
If succeeds -> status="upstreamed"
8. CONFLICT: apply with --force, capture .rej files
patch --force -p1 < {patch_file}
cp *.rej $WORK_DIR/rejections/{patch_name}/
git checkout -- .
status="conflict"
Output per patch:
{
"patch": "0025-clamAV.patch",
"package": "pulpcore",
"status": "conflict|clean|skip|upstreamed|error",
"rej_files": {"pulpcore/content/handler.py.rej": "..."},
"new_file_content": {"pulpcore/content/handler.py": "..."},
"upstream_diff": "..."
}
Status meanings:
| Status |
Meaning |
Who handles it |
skip |
Package not upgraded |
Nobody |
clean |
Applies to new tag as-is |
Nobody |
upstreamed |
Reverse-apply succeeds |
Script deletes; LLM confirms |
conflict |
Hunks failed |
LLM reads .rej files and fixes |
error |
Pre-existing problem |
LLM reports |
04-install-and-apply.sh
INPUTS:
$WORK_DIR/patch-report.json
$PATCHES_DIR
$REQUIREMENTS_FILE
$DEV_CONTAINER_HOST, $DEV_TOKEN
OUTPUTS:
$WORK_DIR/apply-results.json
EXIT CODES:
0 All surviving patches applied
3 One or more patches failed to apply
1 Fatal error (container unreachable, pip failed)
Steps: wait for dev container health, pip uninstall all patched packages, pip
install from requirements.txt, apply each surviving patch in Dockerfile order,
record per-patch result.
Exit code 3 catches patch interaction problems: patch A and B both modify
handler.py, both pass --dry-run against a clean upstream checkout, but B
fails after A is already applied. The report identifies the failing patch and
which prior patch touched the same file.
05-migrate-and-test.sh
INPUTS:
$DEV_CONTAINER_HOST, $DEV_TOKEN
--test-only {spec} (optional, run subset)
OUTPUTS:
$WORK_DIR/migration-log.txt
$WORK_DIR/test-results.txt
EXIT CODES:
0 Migrations and tests pass
4 Migration failed
5 Tests failed
1 Fatal error
Steps: pulp-restart all, pulpcore-manager migrate --noinput,
pulp-test --generate-bindings (or targeted spec).
06-sync-dockerfile.sh
INPUTS:
$WORK_DIR/patch-report.json
$DOCKERFILE
$PATCHES_DIR
OUTPUTS:
Modified Dockerfile (in place)
$WORK_DIR/dockerfile-diff.txt
EXIT CODES:
0 Synced
1 Error
For upstreamed patches (deleted), removes COPY+RUN lines. Verifies every
.patch on disk has Dockerfile entries and vice versa.
07-commit-and-push.sh
INPUTS:
$BRANCH
$WORK_DIR/upgrade-plan.json
$WORK_DIR/patch-report.json
OUTPUTS:
$WORK_DIR/commit-sha.txt
EXIT CODES:
0 Pushed
1 Error
Generates commit message from JSON artifacts:
Upgrade pulpcore 3.108.0 -> 3.109.0, pulp-python 3.28.2 -> 3.29.0
Patches removed (upstreamed):
- 0039-Turn-migration-19-into-a-noop.patch
Patches regenerated:
- 0025-clamAV.patch (conflict in pulpcore/content/handler.py)
Patches unchanged:
- 0010-Added-ability-to-return-a-URL-for-a-blob.patch (oras not upgraded)
Shared library: lib/common.sh
Provides:
dev_exec "command" [timeout] -- runs a command in the dev container via
HTTP API, returns stdout, exits non-zero on failure
log "message" -- timestamped log output
check_artifact "file" -- verifies a JSON artifact exists and is non-empty
Shared library: lib/pkg-workspace-map.sh
Maps pip package names to workspace paths:
declare -A PKG_WORKSPACE=(
[pulpcore]="/workspace/pulpcore"
[pulp-container]="/workspace/pulp_container"
[pulp-python]="/workspace/pulp_python"
[pulp-rpm]="/workspace/pulp_rpm"
[pulp-maven]="/workspace/pulp_maven"
[pulp-npm]="/workspace/pulp_npm"
[pulp-gem]="/workspace/pulp_gem"
[pulp-hugging-face]="/workspace/pulp_hugging_face"
)
# Packages NOT in pulp-stack repo group (clone on demand if upgrading)
declare -A PKG_EXTERNAL_REPO=(
[pulp-file]="https://github.com/pulp/pulp_file"
[oras]="https://github.com/oras-project/oras-py"
[django-storages]="https://github.com/jschneier/django-storages"
)
Goal-Loop Structure
The agent follows a goal-loop, not a linear procedure. Each goal has an
action, a verification condition, and a retry budget.
Goal 1: Versions resolved
- Action: Run
01-resolve-versions.sh
- Verify: exit 0 +
upgrade-plan.json has >= 1 entry in upgrades[]
- On fail: exit 2 = no upgrades, stop cleanly. exit 1 = error, stop.
- Retries: 0
Goal 2: Requirements updated
- Action: Run
02-update-requirements.sh
- Verify: exit 0 +
requirements-diff.txt is non-empty
- On fail: stop
- Retries: 0
Goal 3: All patches resolved
- Action: Run
03-analyze-patches.sh
- Verify:
patch-report.json exists with an entry for every Dockerfile patch
- On fail: exit 1 = fatal error, jump to Goal 7
- Retries: 0 for the script itself
Sub-goal 3a: Upstreamed patches confirmed
For each patch with status="upstreamed":
- Action: Read the patch + new source file at new tag
- Verify: Every change in the patch is present in the new source
- On fail: Restore the patch file, re-classify as "conflict"
- Retries: 0 (read-only check)
Sub-goal 3b: Conflict patches fixed (THE LLM'S CORE JOB)
For each patch with status="conflict":
- Action:
- Read
.rej files from $WORK_DIR/rejections/{patch}/
- Read the new source file(s) at the new tag
- Read the original patch to understand its intent
- Edit the patch file to fix the failing hunks
- Verify:
cd /workspace/{repo} && git checkout {new_tag} && patch --dry-run -p1 < {fixed_patch} exits 0
- On fail: Read the new
.rej output, try again
- Retries: 3 per patch, then mark FAILED and continue to next patch
Goal 4: Packages installed and patches applied
- Action: Run
04-install-and-apply.sh
- Verify: exit 0 +
apply-results.json shows all patches applied
- On fail: exit 3 = read report, find failing patch, fix it (same as Goal 3b), re-run. exit 1 = fatal, jump to Goal 7.
- Retries: 2
Goal 5: Migrations and tests pass
- Action: Run
05-migrate-and-test.sh
- Verify: exit 0
- On fail: exit 4 (migration) = read
migration-log.txt, diagnose, fix, re-run. exit 5 (tests) = read test-results.txt, diagnose, fix, re-run.
- Retries: 2 for migrations, 3 for tests
Goal 6: Dockerfile consistent
- Action: Run
06-sync-dockerfile.sh
- Verify: exit 0
- On fail: jump to Goal 7 (rare failure)
- Retries: 0
Goal 7: Committed and pushed (ALWAYS REACHED)
- Action: Run
07-commit-and-push.sh
- Verify: exit 0 +
commit-sha.txt contains 40-char SHA
- On fail: report error
- Retries: 1
Timeout safety valve
At 80% of the timeout budget (5760s of 7200s), the agent abandons the current
goal and jumps directly to Goal 7. The commit message includes which goals
were completed and which were abandoned.
Progress Tracking
The agent maintains $WORK_DIR/progress.json throughout the run:
{
"started_at": "2026-05-01T08:00:00Z",
"packages": {
"pulpcore": {"current": "3.108.0", "target": "3.109.0"}
},
"patches": {
"0025-clamAV.patch": {
"package": "pulpcore",
"status": "conflict",
"attempts": 2,
"last_error": "hunk 2 FAILED in pulpcore/content/handler.py"
}
},
"goals_completed": ["01-resolve", "02-requirements", "03-analyze"],
"goals_remaining": ["04-install", "05-test", "06-dockerfile", "07-commit"]
}
This enables:
- Observability: A human or the pipeline can inspect this file to see
exactly what happened and why.
- Timeout safety: The agent checks elapsed time against the budget before
starting each goal.
Error Handling
Retry budget
| Step |
Max retries |
On exhaustion |
| 01-resolve |
0 |
Stop |
| 02-requirements |
0 |
Stop |
| 03-analyze |
0 |
Report; conflicts go to LLM |
| LLM patch fix |
3 per patch |
Mark FAILED, continue |
| 04-install |
2 |
Jump to Goal 7 |
| 05-migrate |
2 |
Jump to Goal 7 |
| 05-test |
3 |
Jump to Goal 7 |
| 06-dockerfile |
0 |
Jump to Goal 7 |
| 07-commit |
1 |
Stop |
Always-commit rule
After exhausting retries on any goal, the agent runs 07-commit-and-push.sh
regardless. The commit message includes which packages were upgraded, which
patches were resolved, which failed and why, and which tests failed. The
pipeline creates a draft PR so partial work is always reviewable.
Patch interaction handling
The most subtle error: patch A and B both modify file.py. Both pass
--dry-run against the clean upstream tag. But when applied sequentially in
the dev container, B fails because A already changed the file.
04-install-and-apply.sh detects this (exit code 3) and its report identifies
which patch failed and which prior patch modified the same file. The LLM then
has a precise problem: "patch B fails after patch A is applied to file.py;
here is the file content after patch A; fix patch B to apply on top of it."
Agent YAML
name: bmbouter Upgrade Pulp Dependencies
description: |
Deterministic dependency upgrade agent. Uses shell scripts for all
mechanical work (version resolution, patch testing, package installation,
Dockerfile sync) and reserves LLM reasoning for patch conflict resolution
and test failure diagnosis.
Pass PACKAGES as JSON to target specific versions, e.g.:
{"pulpcore": "3.109.0", "pulp-python": "3.29.0"}
If omitted, discovers and upgrades to latest from PyPI.
prompt: |
You are a dependency upgrade agent. Read
/workspace/pulp-service/.alcove/agents/bmbouter-upgrade-deps/AGENT.md
for your full instructions. Follow the goal-loop structure exactly.
Push your work to the branch specified in the Workflow Context.
Do NOT create a PR -- the pipeline handles that automatically.
timeout: 7200
enforcement_mode: monitor
schedule:
enabled: false
repo_group: pulp-stack
dev_container:
image: ghcr.io/pulp/hosted-pulp-dev-env:main
Pipeline YAML
name: bmbouter Pulp Dependency Upgrade Pipeline
workflow:
- id: upgrade-dependencies
type: agent
agent: bmbouter Upgrade Pulp Dependencies
inputs:
branch: "upgrade/bmbouter-deps"
outputs: [packages_upgraded, patches_modified, patches_removed,
tests_passed, partial_failure]
- id: create-pr
type: bridge
action: create-merge-request
depends: "upgrade-dependencies.Succeeded"
inputs:
repo: pulp/pulp-service
branch: "{{steps.upgrade-dependencies.inputs.branch}}"
title: "Upgrade pulp dependencies"
base: main
draft: true
- id: await-ci
type: bridge
action: await-checks
depends: "create-pr.Succeeded || ci-fix.Succeeded"
max_iterations: 4
inputs:
repo: pulp/pulp-service
pr: "{{steps.create-pr.outputs.pr_number}}"
- id: ci-fix
type: agent
agent: bmbouter Upgrade Pulp Dependencies
depends: "await-ci.Failed"
max_iterations: 3
inputs:
branch: "{{steps.upgrade-dependencies.inputs.branch}}"
ci_logs: "{{steps.await-ci.outputs.failure_logs}}"
outputs: [summary]
- id: code-review
type: agent
agent: PR Reviewer
depends: "await-ci.Succeeded"
inputs:
pr: "{{steps.create-pr.outputs.pr_number}}"
outputs: [approved, comments]
File Layout
.alcove/
agents/
bmbouter-upgrade-deps.yml
bmbouter-upgrade-deps/
AGENT.md
scripts/
01-resolve-versions.sh
02-update-requirements.sh
03-analyze-patches.sh
04-install-and-apply.sh
05-migrate-and-test.sh
06-sync-dockerfile.sh
07-commit-and-push.sh
lib/
common.sh
pkg-workspace-map.sh
workflows/
bmbouter-upgrade-deps-pipeline.yml
Before/After Comparison
| Dimension |
Original |
bmbouter version |
| AGENT.md length |
~300 lines |
~80 lines |
| LLM decision points per run |
~40+ |
3 |
| Upstreaming detection |
LLM reads code, makes judgment |
patch -R --dry-run exit code |
| Patch conflict resolution |
"read upstream diff and recreate" |
Script provides .rej files; LLM fixes specific hunks |
| Deterministic steps |
0% (all LLM) |
~90% (scripts) |
| Repo cloning |
Agent clones mid-run |
alcove pre-clones via repo_group: pulp-stack |
| Progress tracking |
None |
progress.json + JSON artifacts per step |
| Partial failure |
Work lost |
Always commits; draft PR with failure notes |
| Testability |
Must run full agent |
Each script testable independently |
bmbouter Upgrade Pulp Dependencies: Design Spec
Problem
The current
upgrade-depsagent has a 300-line AGENT.md that instructs an LLMto follow an 8-phase procedure. It is non-deterministic because:
shell commands, different judgment calls, different outcomes.
agent must read upstream diffs, decide if a patch is "upstreamed," and
manually regenerate failing patches. This is where the agent goes off the
rails.
patches, or produce broken output with no mechanism to detect it.
phases 1-5 is lost.
Design Principles
Architecture
One agent, seven scripts, goal-loop orchestration.
The LLM has exactly 3 jobs:
status=conflict(read.rejfiles, understand intent,rewrite hunks)
status=upstreamedverdicts (quick confirmation read)Everything else is a script: PyPI queries, version comparison, git tag
checkout, patch testing,
patch -R --dry-runupstreaming detection, pipinstall/uninstall, Dockerfile syncing, commit message generation.
Key Innovation:
patch -R --dry-runfor Upstreaming DetectionThe current agent's #1 failure mode is the "is this patch upstreamed?"
decision. The AGENT.md gives 15 lines of subjective instructions ("check each
change individually," "verify those exact lines exist"). Different runs reach
different conclusions about the same patch.
The replacement is mechanical:
git checkout {new_tag} patch -R --dry-run -p1 < {patch_file}If the reverse of the patch applies cleanly (exit 0), the new version already
contains the patch's changes. No LLM judgment needed, just an exit code.
When a patch genuinely conflicts,
patch --forceproduces.rejfiles thatshow exactly which hunks failed and what the expected vs actual context was.
This gives the LLM a constrained problem ("fix this specific hunk against this
new context") instead of the current open-ended "read the upstream diff and
recreate the patch."
Repo Strategy
The agent uses
repo_group: pulp-stack, which tells alcove to pre-clone allupstream repos into
/workspace/before the agent starts:/workspace/pulp-service(the main repo)/workspace/pulpcore/workspace/pulp_python/workspace/pulp_container/workspace/pulp_rpm/workspace/pulp_maven/workspace/pulp_gem/workspace/pulp_npm/workspace/pulp_hugging_faceScripts reference these pre-cloned repos directly. No git cloning during the
run, no network failures during patch analysis.
Edge case:
orasanddjango-storagesare not inpulp-stack. Patchestargeting those packages (0010 for oras, 0031 for django-storages) are handled
by the script cloning on demand only if those packages are being upgraded.
Script Interfaces
01-resolve-versions.sh
Output format:
{ "upgrades": [ {"package": "pulpcore", "current": "3.108.0", "target": "3.109.0", "workspace_path": "/workspace/pulpcore"} ], "unchanged": [ {"package": "pulp-rpm", "current": "3.35.2", "reason": "already latest"} ] }Uses PyPI
.info.version(never release list sorting) andpackaging.version.Versionfor comparison. The package-to-workspace mappingcomes from
lib/pkg-workspace-map.sh.02-update-requirements.sh
Pure
sedreplacements. Verifies each target version appears in the updatedfile.
03-analyze-patches.sh (the workhorse)
Algorithm (fully deterministic):
Output per patch:
{ "patch": "0025-clamAV.patch", "package": "pulpcore", "status": "conflict|clean|skip|upstreamed|error", "rej_files": {"pulpcore/content/handler.py.rej": "..."}, "new_file_content": {"pulpcore/content/handler.py": "..."}, "upstream_diff": "..." }Status meanings:
skipcleanupstreamedconflicterror04-install-and-apply.sh
Steps: wait for dev container health, pip uninstall all patched packages, pip
install from requirements.txt, apply each surviving patch in Dockerfile order,
record per-patch result.
Exit code 3 catches patch interaction problems: patch A and B both modify
handler.py, both pass--dry-runagainst a clean upstream checkout, but Bfails after A is already applied. The report identifies the failing patch and
which prior patch touched the same file.
05-migrate-and-test.sh
Steps:
pulp-restart all,pulpcore-manager migrate --noinput,pulp-test --generate-bindings(or targeted spec).06-sync-dockerfile.sh
For upstreamed patches (deleted), removes COPY+RUN lines. Verifies every
.patchon disk has Dockerfile entries and vice versa.07-commit-and-push.sh
Generates commit message from JSON artifacts:
Shared library: lib/common.sh
Provides:
dev_exec "command" [timeout]-- runs a command in the dev container viaHTTP API, returns stdout, exits non-zero on failure
log "message"-- timestamped log outputcheck_artifact "file"-- verifies a JSON artifact exists and is non-emptyShared library: lib/pkg-workspace-map.sh
Maps pip package names to workspace paths:
Goal-Loop Structure
The agent follows a goal-loop, not a linear procedure. Each goal has an
action, a verification condition, and a retry budget.
Goal 1: Versions resolved
01-resolve-versions.shupgrade-plan.jsonhas >= 1 entry inupgrades[]Goal 2: Requirements updated
02-update-requirements.shrequirements-diff.txtis non-emptyGoal 3: All patches resolved
03-analyze-patches.shpatch-report.jsonexists with an entry for every Dockerfile patchSub-goal 3a: Upstreamed patches confirmed
For each patch with
status="upstreamed":Sub-goal 3b: Conflict patches fixed (THE LLM'S CORE JOB)
For each patch with
status="conflict":.rejfiles from$WORK_DIR/rejections/{patch}/cd /workspace/{repo} && git checkout {new_tag} && patch --dry-run -p1 < {fixed_patch}exits 0.rejoutput, try againGoal 4: Packages installed and patches applied
04-install-and-apply.shapply-results.jsonshows all patches appliedGoal 5: Migrations and tests pass
05-migrate-and-test.shmigration-log.txt, diagnose, fix, re-run. exit 5 (tests) = readtest-results.txt, diagnose, fix, re-run.Goal 6: Dockerfile consistent
06-sync-dockerfile.shGoal 7: Committed and pushed (ALWAYS REACHED)
07-commit-and-push.shcommit-sha.txtcontains 40-char SHATimeout safety valve
At 80% of the timeout budget (5760s of 7200s), the agent abandons the current
goal and jumps directly to Goal 7. The commit message includes which goals
were completed and which were abandoned.
Progress Tracking
The agent maintains
$WORK_DIR/progress.jsonthroughout the run:{ "started_at": "2026-05-01T08:00:00Z", "packages": { "pulpcore": {"current": "3.108.0", "target": "3.109.0"} }, "patches": { "0025-clamAV.patch": { "package": "pulpcore", "status": "conflict", "attempts": 2, "last_error": "hunk 2 FAILED in pulpcore/content/handler.py" } }, "goals_completed": ["01-resolve", "02-requirements", "03-analyze"], "goals_remaining": ["04-install", "05-test", "06-dockerfile", "07-commit"] }This enables:
exactly what happened and why.
starting each goal.
Error Handling
Retry budget
Always-commit rule
After exhausting retries on any goal, the agent runs
07-commit-and-push.shregardless. The commit message includes which packages were upgraded, which
patches were resolved, which failed and why, and which tests failed. The
pipeline creates a draft PR so partial work is always reviewable.
Patch interaction handling
The most subtle error: patch A and B both modify
file.py. Both pass--dry-runagainst the clean upstream tag. But when applied sequentially inthe dev container, B fails because A already changed the file.
04-install-and-apply.shdetects this (exit code 3) and its report identifieswhich patch failed and which prior patch modified the same file. The LLM then
has a precise problem: "patch B fails after patch A is applied to file.py;
here is the file content after patch A; fix patch B to apply on top of it."
Agent YAML
Pipeline YAML
File Layout
Before/After Comparison
patch -R --dry-runexit code.rejfiles; LLM fixes specific hunksrepo_group: pulp-stackprogress.json+ JSON artifacts per step