What changed, and how each version actually scored.
Every start records which version of the model produced it. That means a version’s record is measured only on the starts it really published, so shipping a new model can never quietly improve an old one’s numbers, and a version that went badly keeps its record. The notes below are written by hand; the numbers beside them are not.
| Version | Published | Starts | Typical miss | Inside range |
|---|---|---|---|---|
| v7.5-arsenal-1.0calibration: pf-backtest base 2023-2025 + live 2026-08-02 | 2026-08-14 – 2026-09-11 | 745 | 1.70 K | 86% |
| v7.3-port-1.1calibration: pf-backtest base 2023-2025 + live 2026-08-02 | 2026-08-07 – 2026-08-13 | 175 | 1.80 K | 85% |
| v7.4-home-1.0calibration: pf-backtest base 2023-2025 + live 2026-08-02 | 2026-08-13 – 2026-08-13 | 4 | 1.82 K | 50% |
| v7.3-port-1.0calibration: pf-backtest base 2023-2025 + live 2026-08-02 | 2026-08-03 – 2026-08-06 | 95 | 1.73 K | 80% |
| v7.3-port-1.0calibration map not recorded, these starts predate version pinning (2026-08-03) | 2026-07-09 – 2026-08-02 | 534 | 1.78 K | 85% |
Typical miss is the average gap between projection and actual. Inside range is how often the result landed within its 80% range, which should be about 80%. Cohorts are not the same size, so a small one says less than a large one.
An audit of what the engine actually applies found that the arsenal-versus-lineup adjustment, meant to be a neutral matchup read, was mis-centered: it raised roughly half the slate's strikeout rates by about three percent on average and almost never lowered anyone's. It has been re-centered on the real distribution so it now reads relative matchups without the built-in lift. Two smaller fixes shipped with it: stacked environment signals (umpire, rest, home-field) are no longer truncated by an over-tight bound, and the public ranges are now graded against what a whole-number range actually promises rather than a nominal percentage that whole numbers can never hit exactly. Tagged as a new model version; its record is graded separately from here.
Checked against held-out 2025 games, the shape of our distributions is well calibrated — but the level was running about a tenth of a strikeout hot on settled starts. The nightly self-review now steps a small global correction toward whatever the last two weeks of graded results say, one percent at a time, hard-capped at three percent in either direction, and it walks itself back to neutral when the drift disappears. Same conservative pattern as the existing workload trim: evidence-gated, bounded, fully audited, and visible on every card it touches.
Checked every graded start from 2025 the model had not trained on: home starters strike out more than road starters, about a third of a strikeout on average, and the projection was ignoring it entirely (the effect was uncorrelated with everything the model already used). It cleared both our tests, the plain signal check and a full re-simulation on unseen games, so we added a small home-field adjustment: a touch up for home starts, a touch down for road ones. Small but real, and tagged as a new model version so its record is kept separate and graded honestly from here.
See the record these versions produced on the results page, or replay any past night exactly as it was published.
The red warning tag was firing on six of every ten starts, mostly restating things the projection already priced in, like walk rate and a contact-heavy opponent. Checked against every graded start, those triggers added nothing, while one buried trigger, a likely blowout, was flagging real trouble the number missed. The warning now fires only for hard external stops: pitch limits, quick-hook managers, badly outmatched teams, and genuinely capped workloads. The dropped reasons still appear in each card's explanation; they just no longer wear a siren.
Fitted a correction aimed at the stated chance of clearing a posted line. It scored worse on days it had not seen, so it was not shipped and the existing model stayed in place. Recorded here because a change that failed its test is still a change we made.
A card must agree with itself: the chance of a quiet night has to be exactly one minus the chance of clearing four strikeouts, and the ladder must never rise as the target gets harder. This now runs on every publish. It immediately caught a real defect, where the quiet-night number was being read off the uncalibrated ladder.
The previous recalibration had graded itself on the same starts it was fitted to, which will always look good. Retuning is now scored on held-out recent days first and is discarded unless it beats the model already in place.
The projected range was too narrow for pitchers with little recent data, which made the model look more certain than it had any right to be. The range now widens with the uncertainty in the inputs.