Harden Aurora backups: 35-day PITR, AWS Backup enrollment, and monthly logical dumps - #492
danielbowne wants to merge 2 commits into
Conversation
|
Thanks for this @danielbowne . Let me start digging in! |
|
Thanks for the writeup, @danielbowne — this is a lot more considered than "needs-refinement" implies, and several things I went in skeptical of held up under checking (Object Lock vs. lifecycle transitions, Your five judgment calls
Changes I'd ask forC1 — blocker:
|
|
Separately from the design feedback above — the deploy-safety gates to clear before this applies (these are pre-apply checks, not change requests):
|
|
@danielbowne checking in on this one. Nothing has moved since the review notes on the 29th, and the PR is still a draft, so I want to make sure it isn't waiting on me. Where it stands: C1 is the only blocker (the Are you picking the C1 fix up, or would you rather I take it? Happy either way, I just want to know whether to plan around it. If this is parked behind other work, say so and I'll leave it alone. |
…l-questionnaire export (#529) Backend release train for the code freeze. Approved PRs only, cherry-picked so each contributor's authorship and signature survive. Frontend counterpart is CMS-Enterprise/ztmf-ui#682. Closes #445 Closes #526 ## What is in it | Source | Author | Change | Closes | | --- | --- | --- | --- | | #522 | @danielbowne | Expose `last_seen` on the users list, plus a login event so it means last sign-in | — | | #524 | @danielbowne | Single home for event-action and score-status values | #445 | | #528 | @MackOverflow, commits by @voidspooks | Export the full questionnaire per system rather than only answered rows | #526 | #522 declares no closing issue by design: it is the backend piece, and presentation is deliberately left open on CMS-Enterprise/ztmf-ui#675. Ordering is by approval. #522 and #524 both touch `events.go` and `users.go`, and they merged without conflict because the edits sit in disjoint regions. #528 is disjoint from both. One thing dropped on purpose: #522's tip was an empty `chore(ci): retrigger DEV deploy under a fresh image tag` commit, pushed to force a fresh image tag past the immutable ECR repo. It has no meaning in a batch that gets its own SHA, so it was skipped rather than carried. ## Verification on the combined branch Verified the combination rather than trusting the parts, since three separately green PRs can still interact: - `go build ./...` and `go vet ./...` clean - `go test -short ./...` — 13 packages ok - Full integration suite against a seeded database — 13 packages ok, zero failures, including the three tests that are sometimes date-dependent - **`make generate-openapi` produces no drift**: the committed spec matches the generated one byte-for-byte. Worth doing here specifically because #522 regenerated the spec and #524 changed models independently, so the combination is the first time those two meet. ## Deploy note, and it is a real one **Do not stack deploys on this branch.** #528's own SHA has two dev deployments two minutes apart: the first succeeded, the second failed with `waiting for ECS Service update: timeout while waiting for state to become 'tfSTABLE' (last state: 'tfPENDING', timeout: 20m0s)`. The second run raced the first while ECS was still rolling out. #524 lost a deploy the same way earlier in the day. So this PR was opened ready rather than draft-then-ready, because that transition is what produced the double run. If the deploy here fails on a stabilise timeout, let ECS settle and re-run **one** deploy; do not push again or trigger a second in parallel. ## Merge order **This merges first.** The frontend batch follows at least five minutes later, which is both a courtesy to the prod pipeline and a correctness requirement: #528's export change has to be live before the frontend's select-all makes never-started systems selectable, or a user can select them and download a header-only file. ## What is not in it - #525 and #492 are drafts. - #423 is a dependency bump with no approval. --------- Co-authored-by: Daniel Bowne <daniel.bowne1@cms.hhs.gov> Co-authored-by: Cameron Testerman <11036339+voidspooks@users.noreply.github.com>
|
Picking this back up next sprint. I drafted it a while ago out of concern about our data-loss exposure, and the substance still stands - cleaning it up against current main for a re-release. |
…and monthly dumps Prompted by an unintended survey response edit discovered five weeks later, with no backup old enough to recover from. Three legs, each covering a different discovery lag: - Point-in-time recovery raised to the Aurora maximum of 35 days in prod. Dev and impl stay at 1 day; neither holds data worth recovering at that horizon. - Enrollment in the CMS OIT Daily15_Weekly90 plan via the AWS_Backup tag the plan already selects on. The plans run daily in these accounts against zero tagged resources today, so this is one tag rather than a new backup stack. - A monthly pg_dump to a new Object Lock bucket, driven by EventBridge Scheduler. This is the only copy with no expiry, and the only one that can be read and diffed without standing up a cluster. PITR alone is not sufficient: the incident surfaced at 35 days, which is the Aurora ceiling, so it would have been covered with zero margin and missed entirely at six weeks. Staleness is measured rather than inferred. CloudWatch caps an alarm's total evaluation range at seven days, so a monthly success pulse cannot be alarmed on across a 40-day window. A daily probe publishes the age of the newest dump instead, which also catches an upload reported as successful that never landed. Also closes drift on the cluster: backup and maintenance windows are pinned rather than left to AWS-assigned random values, the parameter group is pinned to the value already in use, and engine_version is ignored so AWS-applied minor patches stop making every plan propose a downgrade. Verified against dev: plan is 20 to add, 2 to change, 0 to destroy, with both changes in-place on the cluster and its instance.
…ter-group version pin The staleness alarm intentionally fires on an empty bucket, the monthly schedule can be weeks out, and the ops-image build can race the terraform apply for the task definition tag - spell out the enablement steps next to the resources so the day-one ALARM is read as the alarm working. Also breadcrumb the aurora-postgresql16 parameter-group name for the next major version upgrade.
c61f01a to
1cccfc4
Compare
Gives the Aurora cluster a three-legged backup strategy, sized to how late an unnoticed bad edit can plausibly surface:
The incident that prompted this surfaced 35 days after the edit - exactly the Aurora PITR ceiling - so PITR alone covers the worst case already observed with zero margin. The fix is three legs with different horizons: prod PITR raised to the 35-day maximum, enrollment in the CMS OIT
Daily15_Weekly90AWS Backup plan (tag-selected, no new backup stack), and a monthlypg_dumpto a versioned, Object-Locked S3 bucket with no expiry. The dump leg is logical on purpose: recovery for a bad row is "stand up a reference copy and diff", never "roll prod back a month", so readability beats restorability. A daily probe measures the age of the newest dump from the bucket itself and alarms past 40 days - measured, not inferred from the job's own heartbeat, because CloudWatch cannot alarm across a monthly gap.Riders, each argued inline in the diff:
deletion_protection, pinned backup/maintenance windows,copy_tags_to_snapshot, a pinned cluster parameter group, andignore_changesonengine_versionso auto minor upgrades stop producing perpetual downgrade plans.Testing
terraform fmt,tflint, andterraform validatepass on the rebase against current main (clean rebase, no conflicts). The dump script verifies its own artifact (pg_restore --list) and refuses undersized uploads; runtime behavior is exercised in dev first (db_dump_enabled = truethere) before prod depends on it. No application, API, or migration changes - infra and the ops image only.Deploy notes (read before merging)
backend/ops/**and updates theztmf_ops_tagSSM parameter, but the terraform apply in the same pipeline can race it. If the apply wins, the dump/probe task definitions pin the previous image until the next apply. Verify the ops-image run finished, then re-apply (or let the next merge do it).aws ecs run-taskwith the schedule's subnets/SG); the next daily probe reads the fresh dump and the alarm returns to OK. The enablement sequence is also documented at the top ofbackup-dumps.tf.