Automating Baseline Promotion Workflows

A rolling baseline is only as trustworthy as the rule that updates it. If every merge to main overwrites the baseline blindly, the first regression that lands becomes the new normal and the gate quietly stops catching it — the baseline-poisoning failure mode. This guide, part of the Historical Baseline Calibration reference, defines a promotion step that updates the stored baseline automatically, but only after a clean post-merge run on main, and stamps each promotion with the Git SHA that produced it so any bad promotion can be audited and reverted.

The principle is simple: promotion is a privilege, not an event. A run earns the right to become the baseline by passing the existing gate first, and even then only if it is drawn from a healthy window and lands within a sane distance of the value it replaces. Everything below is the mechanics of enforcing that rule in CI. If you have not yet decided which clean event should trigger promotion, pair this with Promoting Baselines After Green Main Builds, which covers the event model in more depth.

Promotion Rules

Promotion is governed by a small set of conditions that must all hold. The table is the authoritative ruleset — a run that fails any row is rejected and the previous baseline is retained.

Condition Requirement Why it gates promotion
Branch Event is a push to main Feature branches carry unmerged, unreviewed code.
Gate status Candidate passed the baseline gate A failing run must never become the reference.
Sample count ≥ 3 runs, median taken A single sample promotes its own noise.
Window health minSamples clean samples after IQR trim Thin windows produce unstable baselines.
Delta sanity New baseline within ±20% of previous A huge jump signals a broken run, not real change.
Provenance Git SHA + timestamp recorded Enables audit and one-command rollback.

The delta-sanity row is the backstop against a broken collection (a blank page, a 500) producing an absurdly fast "improvement" that silently relaxes the band. A change larger than 20% is held for manual review rather than auto-promoted. Read the four hard gates below as an AND, not an OR: the moment any one fails, the run is rejected and yesterday's baseline stands unchanged.

Baseline promotion decision gates A run must clear branch, gate, sample-count, and delta checks in sequence before promotion; any failure retains the previous baseline. Event is a push to main Candidate passed the gate Window has enough clean samples New value within plus or minus 20% Promote median and stamp Git SHA Any failure retains the previous baseline
Promotion is an AND of four checks; a run that misses any one is rejected and the prior baseline is kept unchanged.

Why the Median of a Trimmed Window

Promotion never takes a single number from a single run. It takes the median of a trimmed rolling window, because one Lighthouse run on a shared CI runner routinely swings 10–15% on largest contentful paint even with nothing changing in the code. If you promoted that raw number, you would be laundering runner noise into the reference the whole team is measured against. The safeguard is to keep the last N clean samples, drop outliers with an inter-quartile-range trim, and promote the median of what remains — a construction covered in full in Rolling Median Baseline Windows.

A concrete anchor: if your window holds 30 clean LCP samples with a median of 2150 ms measured on a mid-range mobile profile throttled to Fast 3G, a fresh run that reports 2410 ms should nudge the median only slightly, not redefine the baseline outright. The median's resistance to a single tail sample is exactly why it, and not the mean or the latest reading, is what earns promotion.

Diagnostic Steps

Before automating, confirm what the current baseline is and whether the last main run is eligible to replace it.

# 1. Show the currently stored baseline and its provenance
node scripts/baseline-fetch.js --branch main --print
# → lcp p75=2180ms  sha=4f1c9ad  computedAt=2026-06-19T02:00:00Z  samples=87
# 2. Show the latest main-branch run's median and whether it passed the gate
node scripts/baseline-compare.js --baseline baseline_store.json --run .lighthouseci/ --summary
# → PASS  lcp 2150ms (band 1980-2380)  inp 158ms (band 114-214)  → eligible for promotion

A PASS with a median inside every band, on a main push, is the green light. The LCP band 1980-2380 and INP band 114-214 above are P75 targets for a mid-range mobile profile on Fast 3G; a desktop profile would carry a tighter band. Anything other than PASS and the diagnostic tells you exactly which row of the ruleset blocked the run.

Implementation

The script below promotes the median of a clean run to the baseline. It enforces every rule in the table, recomputes the band from the configured percentile and tolerance, and stamps the Git SHA before writing back to the store.

// scripts/promote-baseline.js  —  run only on a green push to main
import { readFileSync } from "node:fs";
import { execSync } from "node:child_process";

const MIN_SAMPLES = 30;
const MAX_DELTA = 0.20; // delta-sanity guard

function median(xs) {
  const s = [...xs].sort((a, b) => a - b);
  const m = Math.floor(s.length / 2);
  return s.length % 2 ? s[m] : (s[m - 1] + s[m]) / 2;
}

const store = JSON.parse(readFileSync(process.env.STORE_PATH, "utf8"));
const window = JSON.parse(readFileSync(process.env.WINDOW_PATH, "utf8")); // trimmed samples per metric
const gatePassed = process.env.GATE_STATUS === "pass";
const sha = execSync("git rev-parse --short HEAD").toString().trim();

if (!gatePassed) throw new Error("refuse: candidate did not pass the gate");

for (const [metric, cfg] of Object.entries(store.metrics)) {
  const samples = window[metric] ?? [];
  if (samples.length < MIN_SAMPLES) throw new Error(`refuse: ${metric} thin window (${samples.length})`);
  const next = median(samples);
  const prev = cfg.value ?? next;
  if (Math.abs(next - prev) / prev > MAX_DELTA) throw new Error(`refuse: ${metric} delta > 20%`);

  const half = cfg.tolerance.type === "rel" ? next * cfg.tolerance.value : cfg.tolerance.value;
  cfg.value = Math.round(next);
  cfg.band = [Math.round(next - half), Math.round(next + half)];
}

store.provenance = { computedAt: new Date().toISOString(), gitSha: sha };
process.stdout.write(JSON.stringify(store, null, 2));
console.error(`promoted baseline @ ${sha}`);

Pipe the script's stdout back to your store (> baseline_store.json then upload), and let its non-zero exit on any refusal keep the previous baseline in place. Note that the guard compares each metric independently: LCP can be held for a suspicious 30% drop while INP promotes normally, so one broken metric never blocks a legitimate improvement in another.

Guarding Against Baseline Poisoning

Baseline poisoning is what happens when a regression is allowed to define the reference. The moment a slow run overwrites the baseline, the band recenters around the slower number, and every subsequent run — including further regressions — passes against the relaxed band. The gate is still green, but it is measuring against a moving target that only ever moves the wrong way. The chart below contrasts a blind overwrite, which lets a regression at build B4 permanently raise the baseline, with a guarded promotion that rejects the same regression and holds the line.

Guarded promotion versus baseline poisoning A blind overwrite lets a regression become the new baseline and drift upward, while a guarded promotion rejects it and holds the line near its original value. Guarded promotion (rejects the regression) Blind overwrite (poisoned) 2000 2500 3000 regression at B4 2150 ms held 2700 ms B1 B2 B3 B4 B5 B6 B7 Successive merges to main (LCP P75, mid-range mobile on Fast 3G)
The delta and gate checks keep the green line flat near 2150 ms; without them the crimson line drifts to 2700 ms and the gate stops protecting anything.

The lesson is that the delta guard and the gate check are not belt-and-braces redundancy — each catches a distinct failure. The gate check stops a run that is slower than the band from promoting. The delta guard stops a run whose median has moved implausibly far in either direction from promoting, which covers the broken-collection case where a page renders empty and reports an impossibly fast time.

CI Gating Assertion

Wire promotion to run only on main pushes, after the same gate that protects pull requests has already passed. The if: success() on the promotion step is the load-bearing assertion: a failed gate skips promotion entirely.

name: Promote Baseline
on:
  push:
    branches: [main]

jobs:
  gate-and-promote:
    runs-on: ubuntu-latest
    timeout-minutes: 15
    steps:
      - uses: actions/checkout@v4
      - uses: actions/setup-node@v4
        with: { node-version: "20", cache: "npm" }
      - run: npm ci
      - name: Fetch current baseline
        run: node scripts/baseline-fetch.js --branch main --out baseline_store.json
        env: { BASELINE_STORE_URL: "${{ secrets.BASELINE_STORE_URL }}" }
      - name: Collect and gate
        id: gate
        run: |
          npx lhci collect && npx lhci upload --target=filesystem
          node scripts/baseline-compare.js --baseline baseline_store.json --run .lighthouseci/
      - name: Promote (only if gate passed)
        if: success()
        run: |
          GATE_STATUS=pass STORE_PATH=baseline_store.json WINDOW_PATH=trimmed.json \
            node scripts/promote-baseline.js > baseline_next.json
          node scripts/baseline-publish.js --in baseline_next.json --branch main
        env: { BASELINE_STORE_URL: "${{ secrets.BASELINE_STORE_URL }}" }

The separation of jobs on the two triggers is deliberate: pull requests run the identical gate but never reach the promote step, while a push to main runs the gate again against the freshly fetched baseline and only then promotes. The diagram makes the contrast explicit.

PR gating versus main-branch promotion Pull-request runs stop at the gate; only a green push to main proceeds to promote and publish the baseline. Pull request run lhci collect and upload baseline-compare gate PASS allow merge, FAIL block No baseline write on a PR Push to main run lhci collect and upload baseline-compare gate if success, promote median publish store, stamp SHA
The same gate runs on both triggers; only the post-merge run continues past the gate to promote and publish the baseline.

Provenance, Audit, and Rollback

Every promotion writes a provenance record — the short Git SHA and an ISO timestamp — alongside the new value. That record is what turns a bad promotion from a forensic exercise into a one-command reversal. Because each stored baseline is keyed by the SHA that produced it, rolling back means selecting the prior record and republishing it, with no need to recompute anything.

# List the last five promotions with their SHA and value, newest first
node scripts/baseline-history.js --branch main --limit 5 | jq -r '.[] | "\(.computedAt)  \(.gitSha)  lcp=\(.metrics.lcp.value)ms"'
# → 2026-06-20T02:11:00Z  9b2e7c1  lcp=2150ms
# → 2026-06-19T02:00:00Z  4f1c9ad  lcp=2180ms
# Republish the record before a suspect promotion, by its SHA
node scripts/baseline-publish.js --from-history 4f1c9ad --branch main

Because the history is append-only, a rollback is itself a new entry rather than a destructive edit, so you never lose the record of what the poisoned value was. That audit trail also feeds statistical tooling downstream — for example, Changepoint Detection for Performance Time Series can flag the exact merge where a metric shifted, which is often the promotion you want to revert.

Verification

After a merge, confirm the baseline advanced and recorded the new SHA. Re-run the fetch and check that computedAt is fresh and gitSha matches the merge commit.

node scripts/baseline-fetch.js --branch main --print
# → lcp p75=2150ms  sha=9b2e7c1  computedAt=2026-06-20T02:11:00Z  samples=88
git log -1 --format=%h    # → 9b2e7c1  (matches the promoted SHA)

A passing promotion shows the SHA from your latest merge and a computedAt within minutes of it. The lcp p75=2150ms figure is a P75 target for a mid-range mobile profile on Fast 3G; if your fetch prints a value that jumped well beyond the ±20% guard, promotion should have refused and you are looking at either a bypassed guard or a manual override. With promotion automated, the natural next step is catching regressions statistically rather than only against a fixed band; see Automated Regression Detection.

Frequently Asked Questions

Should promotion run on the PR or after merge?

After merge, on a push to main only. A pull request hasn't been reviewed or merged, so promoting from it would let unapproved code define the baseline. Gate the PR with Historical Baseline Calibration, then promote separately on the post-merge run.

What if a legitimate improvement exceeds the 20% delta guard?

The promotion is held, not lost. A real win of that size — say a major code-split landing — should be reviewed and promoted manually by re-running the publish step with the guard relaxed, because a 20%+ swing is far more often a broken run than a real change.

How do I undo a bad promotion?

Re-publish the previous store entry by its recorded Git SHA. Because every promotion stamps gitSha and computedAt, rollback is selecting the prior provenance record and republishing it — no manual recomputation required.

Why promote the median instead of the latest run?

A single Lighthouse run on a shared CI runner can swing 10–15% on LCP with no code change, so the latest reading is mostly noise. Promoting the median of a trimmed rolling window resists that noise; see Rolling Median Baseline Windows for how the window is maintained.

What happens if one metric is broken but the others look fine?

The guard evaluates each metric independently, so a suspicious LCP drop can be refused while INP promotes normally. One broken metric never blocks a genuine improvement in another, and the refusal message names exactly which metric and rule stopped it.