FORECAST FORKThe instability layer of Weather Trader
FF

Forecast Fork

Is the forecast stable, or forking?live from weathertrader.com
A measure of forecast instability

Is the forecast stable,
or forking?

Most weather products tell you the forecast. Forecast Fork tells you whether to trust it — which competing scenarios are still on the table, whether each one is gaining or losing probability run over run, and what it would be worth if it won. It is the open, continuously-verified instability layer of Weather Trader — and every number below is pulled from it live, as you load this page.

See the scenarios live → How it's verified

  The Fork, live & verified in the open

Spread predicts skill
correlation of issue-time ensemble spread with realized Day-5 ACC — negative is the whole point
Runs scored against truth
every forecast verified against what actually happened, in the open
P(bust) model · walk-forward AUC
a fitted bust probability, scored past-only — the deployment-honest metric
The latest run is forking toward
the pattern the current cycle is aiming at
Issue-time ensemble spread  →  realized Day-5 anomaly correlation · the whole archive● live

loading live from weathertrader.com…

This is not a mock-up. It is fetched, as you load the page, from Weather Trader's verification engine: bin every run by the ensemble spread it showed at issue time, plot the skill it actually verified. The down-slope is the proof that spread is a bust signal — known before the outcome. That is the Forecast Fork.

  The fork, drawn — every member counted

A fork is not a metaphor here. Every member of every ensemble — ECMWF EPS, GEFS, AIFS-ENS and all 64 of Google's WeatherNext 2 — is classified into the same named regimes at every lead, and the shares below are those counts, live from the newest cycle. Where the lines split, the atmosphere's futures split.

P(Atlantic regime) by lead · every ensemble member, newest cycle● live

loading the member distribution…

Read it left to right: near the start the members agree and one regime holds all the probability; downstream the shares divide. The lead where they cross is the fork — and it is known, member by member, the moment the cycle is issued.

  The branching graph — scenarios that keep their name

Anyone can cluster an ensemble once. The hard part — and the whole point — is that a scenario must still be the same scenario next run. Every cycle, Weather Trader clusters the pooled members into competing scenarios, then matches this run's clusters to the last run's by area-weighted pattern correlation with an exact one-to-one assignment. Storylines therefore persist: they are born, they gain or lose probability mass, and they die.

Scenario probability, run over run · the Super Ensemble, Days 6–10, N. America● live

loading the scenario tracks…

These are the engine's own composites, rendered on demand — the ensemble mean is the no-clustering baseline, and each scenario is what the atmosphere looks like if that branch wins. Pick your own pool, region, lead and number of scenarios — or build a custom pool from any set of shipped ensembles — on weathertrader.com/clusters.

  The cascade — what a branch is actually worth

A scenario is only interesting if it changes something you can act on. Each branch above is carried through to a downstream quantity: population-weighted degree days over the Lower 48, accumulated across the five-day window, every day scored against its own climatology. Same members, same window, one number per future — plus the probability-weighted expectation and the spread the branches imply.

loading the scenario book…

This is the Fork made spendable. A forecast that is "uncertain" is a sentence; a forecast whose branches differ by cooling degree days over a five-day window is a position size. The full book — four pools, three windows, every k, with each scenario's biggest station movers — is on weathertrader.com/energy.

  Eight models. How many opinions?

A consensus is only worth what its independence is worth. We measure it directly: correlate every model's error field against the same analysis, and count the effective number of independent forecasts. Eight models that all miss the same way are one model with error bars.

Effective independent models · NH Day 5
out of the eight scored against the same analysis
And by Day 10
the fleet herds further onto one error structure as lead grows
Tightest pair
the two models least worth counting twice
Pairwise error correlation ρ(em, en) · NH, Day 5 — the most redundant pairs● live

loading the diversity matrix…

This is why the scenario engine reports each branch's membership by system. A scenario carried by one model family is a different object from the same scenario carried by all of them — and an A.I. model that echoes its IFS/ERA5 training is not an independent vote. Full board: weathertrader.com/herding.

Being precise: the Fork is three quantities

A forecast can be confident — every model and member converging on one outcome — or forked: split into distinct scenarios the atmosphere hasn't chosen between yet. The second kind busts. We keep the three ways of saying that separate, and label each with what has actually been proven about it.

FFspread live · verifiedcurrent uncertainty

Normalized ensemble disagreement per region × lead, fully known at issue time. Its association with realized skill is the chart at the top of this page — measured, not asserted.

FFbust experimentalcalibrated P(fail)

A fitted logistic P(ACC < threshold), scored two ways: leave-init-out, and strict walk-forward — past-only, the honest deployment metric. Reliability is published with it.

FFrevision liveP(the forecast moves)

The Revision Tensor: run-over-run RMS change of the 500 hPa outlook for the same valid day, by lead and region. Temporal instability — distinct from spread, and from P(bust).

FFI = P( materially distinct forecast scenarios remain unresolved )  ·  from cluster divergence, spread growth, cross-model disagreement, run-to-run momentum & bimodality

We combine the three into one headline index only once each is independently, out-of-sample calibrated — not before. Method & live component status on weathertrader.com/method.

The omniscient bookmaker — how good could the consensus have been?

Treat the models as bettors, each staking its Day-5 field. A bookmaker with hindsight — who knows the verifying analysis — would set each run's weights to the blend that minimises the error that actually happened. That book cannot be run in real time. But it can be computed afterwards, and the distance between it and the consensus we actually publish is the information left on the table.

Mean Day-5 blend error (m) · same runs, same truth, four books● live

loading the hindsight ledger…

What the omniscient bookmaker pays each bettor  vs  what our book pays

The method — a branching probability graph

Forecast Fork is not ensemble spread. For a fixed valid time the ensemble is a mixture of coherent scenarios, and we track those branches across successive model runs — not re-clustering and forgetting each cycle. The object is a branching graph: how forecast alternatives emerge, split, merge, and gain or lose probability as the forecast evolves toward reality. The branching graph above is that equation, running.

Ft(x | v) = Σk πk,t · Gk,t(x)  ·  persistent branches · probability-mass flow Δπ · cross-model support · calibrated revision & bust

The system runs in both directions. Forward Fork: from the current state, which outcomes are possible and how are their probabilities evolving? Reverse Fork: given an outcome — the analysis that verified, or a station's daily high above a threshold — which branches carry it, what must stay true, and what would kill it? The branch shares at issue are the prior, each branch's fit to the verifying analysis is the evidence, and Bayes gives the posterior: P(Bk|E) ∝ πk·qk(E) — live on weathertrader.com/reverse.

The Fork family

Fork State livescenario structure

The current branches — how many, how separated, their probability, and their membership by system. Four pools (EPS · full physics · A.I. · the Super Ensemble), five regions, seven leads and windows, k = 2–4, every combination published each run. /clusters

Fork Lineage liveidentity across runs

Scenarios matched cycle to cycle by pattern correlation with exact assignment — births, deaths, and the run-over-run share trajectory of every storyline. The branching graph itself.

Fork Momentum liveprobability-mass flow

Which branch is gaining, losing, stabilizing or reversing — Δπ read straight off the tracks, not re-derived each cycle from scratch.

Fork Cascade liveone branch → all targets

Every branch propagated to a downstream outcome: accumulated HDD/CDD on the census population weights, the stations that move most, the probability-weighted expectation and the spread the branches imply. /energy

Fork Independence livereal vs echoed consensus

The effective number of independent model families — from the pairwise correlation of error fields, plus each scenario's system composition. A.I. models echo their IFS/ERA5 training; this says by how much. /herding

Reverse Fork liveoutcome → pathways

Given an outcome, the branches that carried it: prior share at issue, the evidence, and the Bayes posterior — with the trail of how that branch's probability evolved before it won. /reverse

Fork Revision livethe outlook's own motion

The Revision Tensor — run-over-run RMS change of the outlook for the same valid day, by lead and region, with the current cycle's read against the archive's normal. /verify

Fork Bust experimentalcalibrated P(fail)

A fitted, out-of-sample-scored probability the forecast fails — leave-init-out and strict walk-forward, with the reliability table published. Not merely spread, and not yet promoted.

Fork Autopsy experimentalpost-mortem

After verification: when the winning branch first appeared, how its probability evolved, and what was knowable when. Written up weekly in the Casebook from the frozen record.

Status is honest: live is running publicly and refreshed every cycle, experimental is built and being out-of-sample verified. All of it runs on our own multi-terabyte ensemble archive, computed near the data — member fields reduced on the machines that hold them, so a pooled clustering of two hundred members is a few megabytes on the wire, not a few terabytes. Full method and component status: weathertrader.com/method.

The components — live now on Weather Trader

Engine liveseven systems, one currency

Member fields from every ensemble reduced to a common grid, so EPS, GEFS, AIFS-ENS, WeatherNext 2, HRES, AIFS and GFS can be pooled and clustered as one population.

Index livethe fork number

One number per region & lead — how forked the forecast is — on the Forecast desk's agreement heat.

Scenarios livepooled clusters

Competing patterns cut from the pooled members, each with its own 500 hPa composite, temperature anomaly and degree-day consequence. Build your own on the Cluster Maker.

Calibration liveproven, not claimed

Spread↔skill correlation and a walk-forward-scored bust model, published and continuously updated — the charts above. See the Method.

Attribution livewhy it forks

The synoptic regime the flow is in — blocked, zonal, PNA− — and which model owns it. The named reason a forecast is unstable.

Record livethe standing leaderboard

Mean anomaly correlation at Day 5/7/10 over every run that verified in the last week, month or half-year — all models against one common analysis. /verify

Archive liveissue-time record

Tamper-evident, content-addressed record of every forecast as issued — the sha256 is the identity, so any change is detectable. Hash-chained + externally-anchored roots are on the roadmap.

Why it's different

Not "will it be hot"

  • the full outcome distribution, not a single number
  • the competing scenarios, named and tracked run over run
  • what each branch would be worth if it won
  • whether the forecast is likely to change
  • where and when the models will most likely bust

Public & provable

  • continuously, out-of-sample verified — in the open
  • model-by-model reliability and error, published
  • the consensus's real independence, measured
  • honest probabilities (recalibrated, validated)
  • a recognizable measure anyone can cite

Forecast Fork is the transparent, auditable layer of Weather Trader — a prediction market for the atmosphere. The maps and verification are free and public; the goal is for "the Forecast Fork is high" to become how people say a forecast is unsettled. Start on the live scenario engine or read the method.