Mountain Weather verification programme

How well does the forecast match the hill?

A useful mountain forecast needs more than a convincing explanation. It needs comparison with what was measured, with the strengths and the gaps visible together.

Here is the evidence we can publish today. The next step is to test forecasts saved before the weather happened, separately for each forecast lead time.

Published evidence · development replay

Cairn Gorm summit wind.

A 4.7 mph average mean-wind error across 642 matched sample-hours. This is a useful development check, not a UK-wide forecast accuracy claim.

Location
Cairn Gorm · 1,245 m
Period
7 July to 5 August 2026
Coverage
29 usable dates
Rechecked
5 September 2026
Calculation minus observation, at mast-equivalent height
MeasurementAverage error size (MAE)Average signed error (bias)Correlation
Mean wind4.7 mph-0.18 mph0.90
Gust6.0 mph+1.17 mph0.90

MAE means mean absolute error. It is the average size of the error, ignoring whether it was high or low. Bias keeps the sign: negative means underestimated, positive means overestimated. Opposing errors can cancel in the bias.

Correlation measures the relationship, not percentage accuracy. A value of 0.90 means the calculated and observed winds tended to rise and fall together. It does not mean that 90% of forecasts were correct.

Measurement basis and exclusions

The summit calculation was replayed using retained weather inputs. The comparison is before the walker-height conversion. Both station samples within an hour must pass quality checks, and conflicting duplicate timestamps are excluded. These are sampled conditions within each hour, not continuous full-hour means or peak gusts.

The period was used during development, so it is not an independent holdout. The earlier 658-hour figure used different filtering and is not a paired baseline for claiming improvement.

Sources: Cairn Gorm station information, our summit-wind evidence, and the machine-readable aggregate results. This download contains summary statistics, not station records.

Evidence coverage

Different questions need different tests.

Published development check

Summit wind and gusts

The Cairn Gorm replay above tests the calculation against station samples. It does not give same-day or next-day forecast skill.

Results not yet published

Forecast lead times

No as-issued accuracy results are published here for same day, next day, day 2, day 3 or days 4 to 5. Retained inputs and rebuilt reports cannot fill that gap.

Results not yet published

Temperature, rain and cloud

We are not claiming measured accuracy for these outputs. Each needs compatible observations, matching measurement periods and documented quality checks.

Separate exploratory research

Weather grades and hill activity

The Glen Lyon comparison links historical rain-and-wind categories with recorded hill activity. It is not forecast verification or a test of the complete current hill-day score.

Read the participation findings →

Observation sources reviewed · September 2026

More than one summit. More than one measurement.

We have reviewed the documentation for the sources below to identify ways of testing wind, temperature, rain and cloud across more environments. This is an observation shortlist, not a list of completed validation studies. New accuracy results will appear only after the observations and forecasts have been matched and checked.

UK station archive

Wind, temperature, rain and sunshine

Met Office MIDAS Open through CEDA provides historical station observations under the Open Government Licence. Its separate datasets cover wind, temperature, rain, sunshine and weather observations. Releases extend to the previous complete year, so this is not a live feed.

Met Office Weather DataHub observations are a separate option for current measurements, subject to station coverage, available fields and the chosen access plan.

UK land-surface network

Conditions at more upland sites

UKCEH COSMOS-UK, 2013 to 2025 includes half-hourly temperature, wind, rain, humidity and radiation, with quality flags. This gives us a route to broader land-surface testing, including suitable upland locations.

We need to select sites by elevation and exposure. A sheltered instrument cannot stand in for an exposed summit, and solar radiation is not the same measurement as sunshine duration.

Cloud and visibility

Measure the cloud, then match the mountain

EUMETNET E-PROFILE ceilometers measure cloud-base height and vertical structure. Instrument access and reuse terms need checking site by site. NOAA's Integrated Surface Database also contains cloud and visibility reports, alongside surface weather, where stations report them.

Neither a valley cloud base nor an airport ceiling proves whether a nearby summit was clear. Dated, permission-cleared summit imagery can test clear or obscured conditions separately, not provide an exact cloud-base measurement.

Other mountain environments and how we avoid overstating the evidence

MeteoSwiss automatic stations offer another documented route to wind, temperature, precipitation and sunshine observations. Alpine stations can help test physical behaviour in other terrain, but a Swiss comparison would not establish accuracy in the Scottish hills.

Some station observations appear in more than one archive. MIDAS Open, international surface reports and national rainfall networks are not automatically independent samples. We will identify the original station and measurement, rather than count duplicate records as extra evidence.

Before using a source we need its licence, station coordinates and elevation, sensor height, units, measurement interval, quality flags and retrieval date. Reanalysis is useful for development and climate context, but is not an independent observation of our forecast.

Technical approach · proposed protocol

Keep the forecast. Then check the outcome.

  1. Preserve what was issued

    Save original values, wording, confidence, issue time, valid period and calculation version. A later correction should be a new record, not an edit to the original.

  2. Match like with like

    Check station elevation, sensor height, exposure, averaging interval, gust definition, units and quality flags. Cloud base above the station is not automatically the same as cloud height above sea level.

  3. Separate the tests

    Report lead times, stations, elevation bands, seasons and calculation versions separately. Keep calibration data out of independent evaluation and disclose missing coverage.

  4. Publish misses as well as averages

    Show errors and bias alongside strong-wind hits, misses and false alarms. Report sample sizes and uncertainty, without choosing only the most favourable locations or dates.

What the technical evaluation needs to include

Timing: a report build time or input-cache time is not necessarily the original model issue. A carried-forward forecast keeps its original issue and calculation identity. Unknown provenance stays unknown.

Weather variables: wind and temperature need absolute error and bias. Rain needs defined accumulation periods and event thresholds. A genuine probability forecast can be checked for calibration; an uncalibrated model-support index must not be relabelled as probability.

Cloud: summit-in-cloud and clear-summit observations need their own definitions and quality controls. A valley instrument alone cannot verify whether an adjacent summit was clear.

Fair comparisons: test a simple reference such as persistence where appropriate, using the same cases and information cutoff. Report missing cases, event counts and uncertainty. Many hours in one storm are not many independent storms.

Resolution: point errors do not tell the whole story when rain or cloud is displaced slightly in space or time. ECMWF explains why verification must account for spatial scale. A finer terrain grid is not, by itself, proof of a more accurate forecast.

Progress record

What is available now.

Publication capture and observation shortlist

Added an archive hook to normal peak-forecast publication, preserving wording, values and available version metadata. The first scheduled production capture still needs verification. We have also reviewed further observation sources above. Neither change adds a new accuracy result.

Public verification overview

Published the existing Cairn Gorm development result with bias, average error and its limitations in one place. Set out the proposed as-issued evaluation approach. This is a reporting update, not a new forecast-engine release or a new accuracy result.

Next milestones are verification of the first archived publication, completion of original-issue metadata, observation access and a fixed evaluation protocol. A public archive browser and rolling verification reports are not yet available. Proposed Mountain Weather field stations remain at the planning stage.

We welcome methodological critique and suitable observation-data leads. Named external evaluations will appear only after they have taken place and with permission.

Mountain Weather is operated by Spatial Terrain Limited and led by Simon Grogan, founder, lead developer and mountain leader. About the project.