The reference period
What the anomalies are measured against. The 1991–2020 normal, the period everyone quotes. Reaching it takes a step, because the hindcast spans 2001–2020 and the model's 1991–2020 climatology therefore does not exist — it has to be estimated. It is estimated by shifting the model climatology by the observed change between the two periods, so what is computed is
anomaly(L) = [ forecast(L) − model climatology(L) ] − [ ERA5 1991–2020 − ERA5 2001–2020 ]
The first bracket removes GEPS's drift and its mean bias, because both terms are the same model at the same lead — that is what the hindcast archive is for. The second moves the reference period, and nothing else: it is two decades of climate, measured from observations, with no model in it.
The correction is small and physical, which is the point: 1991–2020 is 0.12 K cooler than 2001–2020 and its 500 hPa surface sits 2.7 m lower, so anomalies quoted here are correspondingly warmer and higher. Precipitation moves 0.03 mm/day, outgoing longwave 0.05 W/m², the winds under 0.06 m/s.
What this deliberately avoids. Differencing straight against ERA5 would also be “vs 1991–2020”, and is the obvious way to do it — but it hands the model's mean bias back into the anomaly, measured here at +0.55 mm/day on precipitation and +9.7 W/m² on outgoing longwave. Those are an order of magnitude larger than the base-period change they would be sitting alongside, and they are a property of the model, not of the weather.
The model climatology and its corrections
The climatology. The GEPS8 reforecast (2001–2020, four starts a week, four members, 39 leads) was pulled from the IRI Data Library for eight variables — 1,920 monthly files, 116 GB — and reduced to a lead-dependent model climatology: for each variable, lead and day-of-year, the mean over hindcast starts within ±15 days of that date across all twenty years, roughly 340–560 starts per cell. Anomalies are taken against that, at matching lead, so what the maps show is departure from how GEPS itself normally behaves that far into a forecast.
Two corrections worth naming, because both were wrong first and both were visible. The climatology is indexed by the day-of-year of the start, so it must be read at the forecast's init day-of-year, the same for every lead; reading it at each valid date instead compared a day-35 forecast against runs launched 35 days later, which over an autumn continent made the reference about 8 K too cold and produced a fake warming that grew with lead. And the ensemble mean must be weighted by member count: the perturbed file holds 20 members and the control one, so averaging the two files equally gave the control half the weight instead of a twenty-first, leaking one member's small-scale noise into every map.
The anomaly is computed on the climatology's grid, not the forecast's. The hindcast is 1° and the live forecast 0.5°, exactly nested. Interpolating the climatology up to 0.5° to keep forecast detail seems free but is not: the forecast resolves terrain and coastlines the 1° reference cannot represent, so that structure does not cancel and comes out as grid-scale speckle along every mountain range and shoreline — mean state masquerading as anomaly, and worst in precipitation. The forecast is now band-limited to the coarse grid first and the difference taken there. An anomaly cannot carry more resolution than its reference.
Outgoing longwave carries a model-version offset, and it is removed. The live operational GEPS radiates about 7.6 W m−2 differently from the GEPS8 reforecast the climatology was built on — flat across all 35 leads (spread 0.53), so it is not a forecast signal; a global-mean OLR anomaly that size is not physically possible. Left in, it tinted every map green and put spurious widespread negative anomalies over North America. The per-lead global mean is removed for OLR, as it is for height. It is irrelevant to the RMM, where a uniform offset moves the projection by less than 0.05.
Height is treated differently from the rest. Between the 2001–2020 hindcast epoch and today there is a real climate offset in geopotential height — a flat +16 m at every lead, measured — which would otherwise sit over every 500 hPa map as a uniform ridge. For z500 the per-lead global mean is removed so the circulation shows. For temperature and precipitation the raw anomaly is kept, because there the departure from the 2001–2020 climate is part of what you want to see. Each figure says which it is.
The OLR channel. ERA5's top-of-atmosphere longwave runs about 13 W m−2 above the NOAA-based reference climatology the RMM EOFs were built on, so a per-longitude offset measured over three whole years is removed before anything is projected. The GMGSI proxy that covers the final few days is offset-corrected onto ERA5 over their 220-day overlap (r = 0.80 in the anomaly, spread within 3% of ERA5's) and screened day by day — the mosaic occasionally publishes a day with missing passes, and one such day at the end of the record would otherwise become the observed MJO's latest point.
The stratosphere
The stratosphere uses ERA5, not the hindcast — because it has to. Everything
else on this page is de-drifted against the GEPS8 reforecast, which is the better
reference. That archive cannot serve the vortex: it carries ua/va
at 100, 200 and 850 hPa only, zg at 200 and 500 only, and no
pressure-level temperature at all. There is no 10 hPa anything in it. So the
stratospheric anomalies are taken against ERA5 instead, and the model bias that the
reforecast would have removed is measured rather than assumed: GEPS analysis against
ERA5 on the same day runs −26 m at 10 hPa globally and +0.5 m over
the polar cap, with temperature within 0.5 K — small enough to leave alone,
and stated here so it is not mistaken for signal. (100 hPa zonal wind is the
one stratospheric field the hindcast could de-drift, and is a candidate for later.)
ERA5 was also checked against MERRA-2, the reference the site’s existing vortex monitor uses, so the two are not silently disagreeing: over 2,156 matching days they differ by 0.13 m/s in u(60°N) at 10 hPa (r 0.999), 0.11 K in cap temperature (r 0.997) and 9 m in cap height from 2001 onward. An archive-wide sweep of both records found a single corrupt day in 46 years of MERRA-2 — 1991-06-30, where the polar-cap height collapses to 0.80 m — now screened out; removing it also moved MERRA-2’s own height trend from +18.7 to +18.1 m/decade, against +18.9 measured independently from ERA5.
Change against the previous run
Change vs previous run. Extended cycles are Monday and Thursday, so the previous run started 3 or 4 days earlier and its forecast for any given calendar day sits at a longer lead. The change panels align the two on the valid date and difference them only where both exist — days 1–32 or 1–31 of the current run — so week 5 is a partial mean and is labelled with the number of common days. Nothing is drawn where only one run has a forecast: zero-padding a partial week would read as “no change” where the truth is “no comparison”. The vortex and teleconnection plots carry the previous run’s ensemble mean as a dashed line on the same rule.
Teleconnection indices
Teleconnections, member by member. The map anomalies are ensemble means; the teleconnection section is the one place on this page built from all 21 members, so that phase can be stated as a probability. The indices are a reproduction of CPC’s, on the reanalysis CPC defines them on (NCEP/NCAR R1, 2.5°), and the reproduction is measured rather than assumed. The AO and AAO are the leading EOF of monthly sea-level pressure poleward of 20° in each hemisphere (Thompson & Wallace 1998), signed so the positive phase has low pressure over the pole — CPC builds its own on 1000 hPa and 700 hPa height respectively, so the AAO in particular is a close relative rather than a copy; projected day by day they track CPC’s daily AO and AAO at r 0.97 and 0.96. The 500 hPa patterns (NAO, PNA, EA, WP, EP/NP, EA/WR, SCA, TNH, POL) are defined by regression: for each pattern and each calendar month, the map of monthly height anomalies 20–90°N — standardized by calendar month and weighted by √cos φ, as in Barnston & Livezey (1987) — regressed on CPC’s published monthly index over the months within one of that calendar month, and the nine maps of a month combined into a least-squares projector so each index is read with the others accounted for. Month-by-month matters: CPC derives its modes separately for every calendar month and the PNA in particular changes shape with season, so a single all-year rotated PCA reproduced CPC’s PNA at only r 0.54, the seasonal regression at 0.74. Fidelity is scored out of sample — patterns fitted on 1991–2005, scored on 2006–2020 against CPC’s monthly indices for every pattern and against its daily NAO and PNA — and those are the correlations in the table; the operational projector is then refitted on all thirty years. A daily index is the daily anomaly projected the same way and divided by the seasonal standard deviation of the daily projection — the spread within ±15 days of that day of year over 1991–2020 — so “+1” means one standard deviation for the time of year. CPC standardizes its daily AO, NAO and PNA by a single annual figure instead; on that convention the indices projected from raw height or pressure (AO, AAO, NAM, EPO, WPO) read as permanently quiet in summer and every ±0.5 threshold is easy in January and hard in July. The published CPC and PSL values drawn on the plumes are rescaled to the seasonal unit, and the correlations quoted against them are computed on the annual convention, like for like. Each GEPS member’s de-drifted, re-based 500 hPa and sea-level-pressure anomaly — the same anomaly the maps show — is projected identically, so the forecast index carries no model drift; the tail is GEPS’s own day-0 analyses of the last weeks, anomalised the same way (against the model climatology at lead 0.5 and re-based), so analysis and forecast sit on one rule — against the NCEP climatology instead, the GEPS-versus-R1 height offset projected onto the monopole-like patterns as a spurious step at the seam. CPC’s own published daily values are drawn beside the tail. Two Pacific indices forecasters use that are not in CPC’s set are added on PSL’s definitions — box differences of 500 hPa height anomaly, EPO (20–35°N minus 55–65°N over 160–125°W) and WPO (25–40°N minus 50–70°N over 140°E–150°W) — and reproduce PSL’s daily series at r 0.97 and 0.99. The NAM at 100 hPa, the stratosphere–troposphere coupling index, is the leading EOF of 100 hPa height anomalies poleward of 20°N with the hemispheric mean removed first (the +19 m/decade height trend at that level would otherwise read as a permanently weak vortex), taken from the 21 members’ 100 hPa heights; the GEPS8 hindcast carries nothing above 200 hPa, so it is a plain anomaly against the NCEP climatology and gets no calibration. Raw probabilities are member fractions, so they resolve to about 5% and 0% means “no member”, not certainty.
Calibrated on the hindcast. The same projectors are applied to every GEPS8 hindcast start of 2001–2020 — four a week, 39 daily leads, de-drifted and re-based exactly as the live product — and scored against the observed index on the valid date from NCEP/NCAR R1. For each index and week, the starts within ±45 days of the current init day-of-year (about a thousand) give the anomaly correlation of the hindcast ensemble mean (the skill quoted in the table) and a regression of the observed weekly index on the forecast one; today’s ensemble-mean weekly index passed through that regression, with the residual spread as the distribution, is the calibrated probability. Where the model has no skill the regression slope goes to zero and the calibrated probability collapses to climatology, which is the honest statement for a week-5 NAO; where it has skill but too little amplitude, the slope restores it. The hindcast mean is four members against 21 live, so the slope is slightly conservative. A weekly mean of a daily-sd index is a smoother quantity than the daily values, so a weekly mean beyond ±1 is a strong signal.
Skill masks and tercile probabilities
Skill mask and probability maps. The same hindcast that calibrates the indices is scored gridpoint by gridpoint: every 2001–2020 start’s weekly-mean anomaly, de-drifted and re-based exactly as the live maps, against an independent analysis on the same 2.5° grid — NCEP/NCAR R1 2 m temperature, 500 hPa height and sea-level pressure, and the CPC unified gauge analysis for precipitation, which exists over land only — pooled by calendar month. The anomaly correlation from that is the skill mask (hatched below 0.25; cross-hatched where there is no truth), and the gridpoint regression of observed on forecast weekly anomaly, with its residual spread, turns the ensemble mean into calibrated tercile probabilities, below, near and above normal, drawn CPC-style as the most likely category with its probability as intensity. The terciles are those of the 2016–2025 weekly anomaly for that week of the year, not 1991–2020: against the older base the 2026 maps read “above” nearly everywhere before the forecast says anything, which is the warming since the base period, not weather. The calibration is applied to the departure from the recent median, using the hindcast slope but not its intercept — the intercept is the 2001–2020 mean state, and a no-skill regression that returned it printed the hindcast-era climate as “below normal” against a recent threshold. With no skill the three probabilities go to a third each. Near-normal is the tercile models forecast worst; its probability rarely clears 40% beyond week 1, and with 21 members its raw fraction is noisy. Precipitation terciles are undefined where the climatological week is dry, so those cells are masked as CPC masks them. The raw member fraction is shown beside the calibrated map: where the two disagree the hindcast is saying the ensemble is over- or under-confident there. Gaussian residuals are an approximation for precipitation; at a weekly mean and 2.5° it is a fair one.
Weather regimes and ensemble scenarios
Weather regimes. The Euro-Atlantic and North American sectors each get four regimes from k-means on 5-day-mean 500 hPa anomalies (NCEP/NCAR R1, 1991–2020, the leading 14 EOFs, 30 restarts), separately for the cold half-year (October–March) and the warm (April–September), because summer regimes are not winter regimes with less amplitude. The cold Euro-Atlantic set comes out as the classic four — NAO+, Scandinavian blocking, Atlantic ridge, NAO−/Greenland blocking — and the North American set as Pacific trough, Alaskan ridge, Pacific ridge and Greenland high, the Lee et al. (2019) quartet; the warm-season sets are named by inspection and say so. Each of the 21 members’ de-drifted 500 hPa anomaly is smoothed the same way and assigned to the nearest centroid, or to “no regime” when its pattern correlation with every centroid is below 0.25, as in the ECMWF product; the member fraction is the regime probability. The set for the init month’s season is held for the whole forecast so a plume crossing the October or April boundary is not reclassified mid-way. The hindcast hit rate below the plume is the share of 2001–2020 starts within ±45 days of the date whose ensemble-mean regime matched the observed one, against persistence of the initial regime and the most common regime. On raw k-means of height anomalies alone the value is modest, and the hit rate shows it: what earns the panel its place is the member fractions with that skill curve beside them.
Ensemble scenarios. Distinct from the regimes: no fixed centroids. Each week the 21 members’ weekly-mean 500 hPa anomalies over a sector are clustered by k-means with k = 2, 3 and 4, and the number of scenarios is chosen, not fixed: each k is tested against a null in which the members’ deviations are rotated by random orthogonal matrices in member space — the spatial covariance is kept exactly, any grouping is destroyed — and re-clustered a hundred times with the same k-means effort. The k that beats the 95th percentile of its null by the widest margin is drawn as scenario means with member counts; if none does, the ensemble is one scenario with spread and the two-way split is shown faded. This is the Ferranti and Corti (2011) recipe ECMWF uses to pick its cluster count. It was calibrated on this ensemble: Gaussian ensembles pass 5–7% of the time, a two-scenario ensemble with groups three standard deviations apart is caught 60% of the time at k = 2 and 95% at four apart, but a forced k = 3 caught those only 25% and 50% of the time, which is why a fixed three-way split was abandoned. Groups closer than about two standard deviations cannot be told from a continuum with 21 members, and the usual verdict is one or two scenarios. Each scenario mean is labelled with its nearest weather regime and the pattern correlation.
MJO skill
MJO skill. Every hindcast start is projected onto the same Wheeler–Hendon EOFs with the same arithmetic as the live RMM — band means, lead-matched de-drift, frame shift, 120-day mean removed — and verified against the Bureau of Meteorology’s RMM on the valid date, giving the bivariate correlation, RMSE and amplitude ratio by lead, for all starts, by season, for strong initial MJO, and for the starts within ±45 days of today. The one approximation: the WH04 120-day mean before each start is taken from the hindcast’s own day-1 channels of the preceding 120 days rather than from observations, which are not in this archive for 2001–2020; day-1 skill is ~0.95 and the term is slowly varying. The ECMWF marks are the published extended-range reforecast values (Vitart 2017 and the ECMWF verification memoranda), not a like-for-like computation — the S2S reforecasts sit behind a research licence and are not on this site, and ECMWF’s open data stops at day 15, so no live comparison beyond that is possible either.
Reading the MJO score honestly. Over all 4,794 starts the bivariate correlation falls through 0.6 at day 15 and 0.5 at day 19; winter starts hold on longest (0.6 at day 16), the late-summer window the site is in now is the weakest (0.6 at day 13). Against the published ECMWF figures of day 27–30 that is about two weeks of skill less, and the following caveats move it a few days, not two weeks. Ensemble size: the hindcast mean is four members and the operational GEPS runs 21; a larger mean removes unpredictable noise and is typically worth two to three days at the 0.6 level, while ECMWF’s published numbers come from an eleven-member mean, closer to their operational system. Model vintage: GEPS8 is the version this reforecast belongs to, not the current cycle, and ECMWF’s quoted range spans 2016 to the early 2020s. Verification target: BoM’s RMM removed an ENSO component before 2014 and not after, and ECMWF scores against its own analysis; a day at most. The 120-day mean: approximated from the hindcast’s own day-1 channels, a slowly varying term with day-1 skill near 0.95. What the scores describe: the ensemble mean only — the plume’s spread is not scored here, and an ensemble mean is damped by construction, which is what the amplitude panel shows. The practical reading is that the MJO phase diagram on this page is reliable to about two weeks, useful with care to about three, and beyond that is a model opinion.
Caveats and data
Caveats. The maps are an ensemble mean, so they understate the range of outcomes everywhere and increasingly so with lead; outside the teleconnection section nothing here is a probability. Twenty-one members is also not many for precipitation: a single member differs from the member mean by about the field's own magnitude, so roughly a fifth of single-member variance survives into the mean and shows up as coherent-looking blobs at long lead. Those are ensemble sampling noise, not forecast structure. Extended cycles run Monday and Thursday only. The hindcast is a fixed model version while the live forecast is the current operational one, so a slow model change is not captured by the 2001–2020 reference — the outgoing-longwave offset above is exactly that showing up.
Data: ECCC GEPS via the MSC Datamart (Environment and Climate Change Canada, Open Government Licence). GEPS8 reforecast via the IRI Data Library, Columbia University. ERA5 via the Analysis-Ready Cloud-Optimized store, Copernicus Climate Change Service. GMGSI via NOAA on AWS. RMM reference EOFs after Wheeler & Hendon (2004).