Skip to content

Commit bb8e046

Browse files
test: recalibrate Windows warm P50 budget (Fixes #513) (#515)
## Summary Recalibrate only the Windows schema-v2 warm full-refresh P50 budget after unchanged product runs exposed an unusually low exact baseline. The new limit accepts the measured 105-182ms range while still blocking a sustained median above 255ms against that baseline. ## Changes - change Windows warm full-refresh P50 from 50ms/30% to 150ms/50% - preserve Linux/macOS, P95, cold, startup, coverage, and schema gates - prove the observed 182ms run passes - prove a 300ms sustained warm median fails without implicating P95 or cold P50 - preserve explicit coverage that both absolute and relative limits must be exceeded - document the six-measurement calibration set ## Validation - `python -m unittest discover -s scripts/tests -p 'test_*.py' -v` (35 passed) - observed 182ms Windows artifact passes against exact baseline `ad7ca14` - all-platform performance run `31524812300` passed after retrying a transient baseline-download certificate failure - clean Copilot review on superseded stacked PR #514 Supersedes #514, which GitHub automatically closed when its stacked base branch was deleted. Fixes #513 --------- Co-authored-by: Copilot <223556219+Copilot@users.noreply.github.com>
1 parent ae216ca commit bb8e046

3 files changed

Lines changed: 25 additions & 5 deletions

File tree

‎docs/QUALITY_SNAPSHOTS.md‎

Lines changed: 4 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -16,18 +16,20 @@ A metric blocks when it exceeds both its absolute and relative budget:
1616
| --- | ---: | ---: | ---: |
1717
| Server startup P50 | 5 ms / 100% | 10 ms / 50% | 100 ms / 50% |
1818
| Server startup P95 | 50 ms / 200% | 50 ms / 100% | 750 ms / 100% |
19-
| Full refresh P50 | 25 ms / 30% | 50 ms / 30% | 100 ms / 50% |
19+
| Full refresh P50 | 25 ms / 30% | 150 ms / 50% | 100 ms / 50% |
2020
| Full refresh P95 | 50 ms / 50% | 250 ms / 100% | 300 ms / 100% |
2121
| Time to first environment P50 | 20 ms / 100% | 25 ms / 50% | 150 ms / 50% |
2222
| Time to first environment P95 | 25 ms / 100% | 100 ms / 100% | 250 ms / 100% |
2323
| Cold refresh P50 | 100 ms / 50% | 150 ms / 50% | 250 ms / 50% |
2424

25-
Each cell is `absolute / relative`. The warm P50 and server-startup budgets reflect observed GitHub-hosted runner variance from 11 consecutive main-branch baselines. Tighten them when a noisy path is fixed rather than normalizing a known regression into the baseline.
25+
Each cell is `absolute / relative`. The Linux/macOS warm P50 and all server-startup budgets reflect observed GitHub-hosted runner variance from 11 consecutive main-branch baselines. Tighten them when a noisy path is fixed rather than normalizing a known regression into the baseline.
2626

2727
The macOS server-startup P95 budget recalibration is tracked by issue #507 and follows PR #506's fix for issue #504. It uses three unchanged-content pull-request runs and the exact merged baseline at `f0c62d9`; the resulting absolute headroom is four to six times the observed post-fix run-to-run range.
2828

2929
The warm refresh and warm time-to-first P95 budgets were recalibrated in issue #511 after PR #510 separated cold and warm samples. Three unchanged-code PR runs plus the exact schema-v2 baseline at `ad7ca14` retain at least 2.5 times the observed absolute run-to-run range.
3030

31+
The Windows warm full-refresh P50 budget was recalibrated in issue #513 from five unchanged-code pull-request measurements plus the exact schema-v2 baseline at `ad7ca14` (six measurements total). It retains nearly twice the observed absolute range while blocking a sustained median above 255ms against that baseline.
32+
3133
Schema v2 records `full_refresh` and `time_to_first_env` from the warm member of each pair and adds cold refresh/time-to-first distributions. During its one-time rollout, comparisons against a schema-v1 base checked cold P50 against explicit absolute ceilings of 500ms on Linux, 750ms on Windows, and 1,000ms on macOS. Schema-v2-to-v2 comparisons use the table's dual budgets.
3234

3335
The cold P50 budgets were calibrated in issue #509 using two unchanged-head all-platform runs and the final pull-request validation.

‎scripts/quality_snapshot.py‎

Lines changed: 1 addition & 1 deletion
Original file line numberDiff line numberDiff line change
@@ -91,7 +91,7 @@ def regressed(self) -> bool:
9191
'windows': (
9292
RegressionBudget(10, 50),
9393
RegressionBudget(50, 100),
94-
RegressionBudget(50, 30),
94+
RegressionBudget(150, 50),
9595
RegressionBudget(250, 100),
9696
RegressionBudget(25, 50),
9797
RegressionBudget(100, 100),

‎scripts/tests/test_quality_snapshot.py‎

Lines changed: 20 additions & 2 deletions
Original file line numberDiff line numberDiff line change
@@ -73,7 +73,7 @@ def test_unchanged_snapshot_passes(self):
7373
self.assertEqual(failures, [])
7474

7575
def test_p50_regression_fails_when_both_budgets_are_exceeded(self):
76-
current = performance_snapshot(refresh_p50=180)
76+
current = performance_snapshot(refresh_p50=300)
7777
_, failures = compare_performance(current, performance_snapshot(refresh_p50=100), 'Windows')
7878
self.assertTrue(any('Full refresh P50' in failure for failure in failures))
7979

@@ -158,6 +158,24 @@ def test_newer_performance_schema_is_invalid(self):
158158
'Windows',
159159
)
160160

161+
def test_schema_v2_windows_warm_p50_variance_passes(self):
162+
baseline = performance_snapshot(schema_version=2, refresh_p50=105)
163+
current = performance_snapshot(schema_version=2, refresh_p50=182)
164+
165+
_, failures = compare_performance(current, baseline, 'Windows')
166+
167+
self.assertEqual(failures, [])
168+
169+
def test_schema_v2_windows_warm_p50_material_regression_fails(self):
170+
baseline = performance_snapshot(schema_version=2, refresh_p50=105)
171+
current = performance_snapshot(schema_version=2, refresh_p50=300)
172+
173+
_, failures = compare_performance(current, baseline, 'Windows')
174+
175+
self.assertTrue(any('Full refresh P50' in failure for failure in failures))
176+
self.assertFalse(any('Full refresh P95' in failure for failure in failures))
177+
self.assertFalse(any('Cold refresh P50' in failure for failure in failures))
178+
161179
def test_schema_v2_warm_p95_variance_passes_on_all_platforms(self):
162180
cases = (
163181
('Linux', 60, 69, 16, 16),
@@ -233,7 +251,7 @@ def test_noise_inside_absolute_budget_passes(self):
233251

234252

235253
def test_relative_budget_must_also_be_exceeded(self):
236-
current = performance_snapshot(refresh_p50=1_060)
254+
current = performance_snapshot(refresh_p50=1_160)
237255
_, failures = compare_performance(current, performance_snapshot(refresh_p50=1_000), 'Windows')
238256
self.assertEqual(failures, [])
239257

0 commit comments

Comments
 (0)