Methodology
What’s measured, what’s modelled, and what’s made up.
Where each number comes from
The Sonde forecast
Three terms, and only one of them is invented.
Model consensus (real). A skill-weighted mean of ECMWF, GFS and ICON for the market’s target — daily maximum or overnight minimum — at the station’s coordinates, in the station’s own timezone. Open-Meteo’s best_match blend carries double weight because it is already a skill-weighted product.
Observation bias correction (real). Today’s actual observations are compared hour-by-hour against what the models said those hours would be. If the station has run 1.2° warm all morning, the afternoon peak is very likely under-forecast too. The mean residual is applied at 65% weight and clipped to ±3°, because part of any early divergence is timing rather than amplitude.
Balloon adjustment (simulated). A bounded term, roughly ±1.8°, derived from the generated boundary-layer profile — deeper mixing and a steeper low-level lapse rate support overshooting model guidance. On a real deployment this is where fresh sounding data would enter. Here it is fabricated, and every API response breaks it out separately so you can subtract it.
The hourly curve is the model’s diurnal shape shifted so its extremum lands on the forecast value. Peak timing is the argmax refined by a parabola through its neighbours.
Uncertainty and fair value
σ starts at 0.35°, grows with hours remaining to the extremum and with the spread between the public models, and collapses once the day’s observed maximum has already reached the forecast — at that point the remaining uncertainty is small and one-sided. Overnight lows carry a flat +0.25° penalty for radiational-cooling uncertainty.
Fair value for a bucket is the normal integral over it. Critically, both venues resolve on the whole-degree climate report, so a “91–92°” bucket means the rounded value lands in {91, 92} — a continuous high anywhere in [90.5, 92.5). Ignoring that half-degree skirt systematically misprices every boundary bucket, which is precisely where the edges are.
Edge is fair minus price on the best tradeable line, considering both sides: YES is priced at the ask, NO at 100 minus the bid. Pricing off the touch rather than the mid is what stops a 20¢ “edge” that is really 20¢ of spread. Sizing is quarter-Kelly, ¼·(q−p)/(1−p).
Market data
A Kalshi event is one market here; the Kalshi markets inside it are the ladder rungs, with bounds read from the structured strike fields. Polymarket has no structured strikes, so ranges are parsed from the outcome titles — an event that doesn’t yield at least two parseable rungs is skipped rather than guessed at, because a half-parsed ladder produces confidently wrong fair values.
Market-implied temperature is the probability-weighted midpoint of the ladder after normalising away the venue’s overround.
Production only shows real venue books. Demo placeholder ladders can be enabled with SONDE_DEMO_MARKETS=fallback when a venue returns no open market for a station. Those rows are badged DEMO BOOK, counted separately in diagnostics, and flagged simulated.market_book in the API.
The track record is not a track record
A real version of this product would archive every forecast at issue time and score it months later. This deployment has no such archive, so the research page reconstructs one: resolved temperatures are real, and each source’s historical “forecast” is drawn from a fixed per-source error distribution, then settled through the same bucket-selection and quarter-Kelly code a real archive would run through.
That makes the shape of the analysis honest — calibration, Brier, drawdowns and the P&L path are all computed for real — while the skill being demonstrated is assumed rather than measured. Do not read the MAE table as evidence that anything beats ECMWF. It is a fixture.
Known limits
Model run times are derived from the nominal 00/06/12/18Z schedule and each model’s publication latency — Open-Meteo doesn’t expose an actual run stamp on the forecast endpoint, so those labels are cycles, not issue times.
Historical “resolved” values come from Open-Meteo’s analysis, not from the station’s official CLI climate report, and can differ by a degree at the rounding boundary.
Volume figures are venue-reported and not normalised across venues: Kalshi contracts settle at $1 so traded notional is contracts × price, while Polymarket reports USDC volume directly.
Not financial advice
This is a demonstration of a data product. The forecast contains a fabricated component by design, the historical performance is a fixture, and nothing here has been validated for trading. Do not size positions off it.