Skip to main content
Quanta Meridian logo
Skip insight types

Research Brief

When is a better forecast strong enough to test further?

The R correction has a lower average error. Why was the published forecast still retained?

Because the fixed test decides whether a challenger deserves further controlled review, not whether one retrospective result should replace the operating forecast. R cleared that test; Python did not; neither result authorised an operating change.

Quanta MeridianPublished Reviewed 7 minute research briefResearch Brief 3 of 3
How six rolling origins lead to a bounded model-selection decision
  1. 01Forecast issueOne published operating reference

    The comparison begins with the forecast available at the recorded issue time.

  2. 02Rolling originsSix later 28-day tests

    Each origin fits the previous 730 days and retains the following outturn.

  3. 03Comparable seriesPublished · R · Python

    The same 8,066 half-hours support every model comparison.

  4. 04Selection gatePoint error, consistency and interval behaviour

    A challenger advances only when every pre-agreed check passes.

  5. 05Operating referencePublished forecast retained

    R earns further controlled review; Python does not clear the complete gate.

Day-ahead electricity demand forecast review comparing the published forecast, R and Python corrections and later outturn across a fixed model-selection gate
Executed rolling-origin evaluationSix origins, three comparable forecasts and one fixed gate

The curve and selection result are generated from the retained published forecast, R and Python predictions and later outturn. All 8,066 held-out rows remain available in the complete example.

The review compares the published day-ahead electricity-demand forecast with an R linear correction and a Python boosted correction across six fixed rolling origins. All three forecasts are scored against the same 8,066 later half-hour outturns. R improves mean absolute error by 2.1% and wins five origins. Python improves it by 0.7%, while its nominal 80% interval covers 72.6% of outcomes. Those results support different next steps, not an automatic model replacement.

The selection rule was fixed before the result

A challenger must improve aggregate mean absolute error by at least 2%, beat the published forecast in at least four of the six origins, avoid a material deterioration in bias or peak-period error, and place between 75% and 85% of later observations inside its nominal 80% interval. Passing every check means advance to controlled review. It does not mean replace the operating reference.

Using one gate prevents the strongest-looking percentage from deciding the outcome after the analysis has been seen. It also keeps point error, consistency and interval behaviour in the same decision.

Six origins show whether the improvement travels

Each origin uses the previous 730 days and predicts the following 28 days. The tests start in January, April, July and October 2025, then January and April 2026. Later outturn stays outside the training data until each forecast is scored.

R records a lower mean absolute error in five origins and improves the aggregate result from 697.9 MW to 683.0 MW. Python wins four origins but improves the aggregate result only to 693.0 MW. The six-origin view therefore shows both where a correction helps and where it loses; it does not present one favourable period as the result.

The Python interval changes the decision

Python reduces aggregate mean absolute error by 0.7% and controls aggregate bias, but it misses two parts of the fixed gate. Its improvement is below the 2% threshold and its nominal 80% interval contains 72.6% of later outcomes, below the accepted 75% to 85% range.

That coverage result matters because a narrow-looking range can understate uncertainty even when the point forecast improves. Python therefore does not clear the complete gate. The comparison does not describe its interval as calibrated.

Further review is not replacement

R passes all five checks: 2.1% lower mean absolute error, five origin wins, lower bias, lower peak-period error and 79.4% empirical coverage. The correct next step is to examine it beside the current forecast under controlled conditions. The published forecast remains the operating reference.

The traced record, 2026-04-09/SP30, reinforces that boundary. The published forecast was 20,708 MW and the later corrected outturn was 25,045 MW, but the record was selected because it was the largest published-forecast miss in the final origin. It is an inspectable error review, not a typical half-hour or evidence that the correction will perform the same way in future.

What the comparison cannot establish

The retained study is retrospective and ends on 30 June 2026. It does not reproduce National Energy System Operator forecasting, run as a live service or include actual weather, embedded generation, special events or live system conditions as issue-time inputs.

Historical error reduction does not prove causal benefit, future accuracy or operational safety. R intervals depend on the fitted linear-model assumptions, and the Python interval remains outside the stated coverage gate. Any further test needs current operational inputs, an accountable forecast owner and a separately approved change decision.

Research summary

Fact, interpretation and implication

The measured result is shown separately from its interpretation and the action a practitioner may consider.
Research question
Does either proposed correction improve the published day-ahead forecast consistently, with acceptable interval behaviour, strongly enough to enter further controlled review?
Source
Quanta Meridian: Day-ahead Electricity Demand Forecast Review; National Energy System Operator: Day Ahead Half Hourly Demand Forecast Performance; National Energy System Operator: NESO Open Data Licence v1.0
Method
The retained build uses six fixed rolling origins. Each fits the previous 730 days and predicts the next 28 days without using later outturn at issue time. Published, R and Python results are compared across the same 8,066 held-out half-hours using mean absolute error, origin wins, bias, peak-period error and empirical interval coverage. The complete gate was declared before the aggregate result was reviewed.
Observed finding
R improves mean absolute error by 2.1%, wins five of six origins and passes all five gate checks. Python improves mean absolute error by 0.7% and wins four origins, but its improvement and 72.6% interval coverage miss the complete gate. R advances only to further controlled review; the published forecast remains the operating reference.
What the finding does not establish
This is a retrospective public-data study, not a live forecasting service or a reproduction of National Energy System Operator forecasting. It excludes actual weather, embedded generation, special events and live operating conditions. Historical improvement does not establish causal benefit, future accuracy or authority to change an operating forecast.
Practical implication
A challenger should earn further testing through a rule fixed before its result is known. Point error, consistency, peak performance and interval behaviour should lead to an accountable review decision; no single lower average should replace an operating reference automatically.

Comparison

Aggregate result and selection decision

Mean absolute error, origin wins and interval coverage are reproduced from the retained model evaluation. Lower error is better.
ForecastMean absolute errorChange from publishedOrigin winsInterval coverageDecision
Published forecast697.9 MWOperating referenceNot applicableNot reportedRetain as operating reference
R linear correction683.0 MW2.1% lower5 of 679.4%Advance to controlled review
Python boosted correction693.0 MW0.7% lower4 of 672.6%Does not clear the complete gate

Related project examples

See how the finding was calculated and checked

The linked projects show the records, calculations and checks used for the finding. Open one for the full project detail.

Sources

Sources used to answer the research question

  1. Day-ahead Electricity Demand Forecast ReviewQuanta Meridian

    Retained Python and R build, six rolling origins, 8,066 held-out predictions, selection gate, settlement-period trace and explicit limitations.

  2. Day Ahead Half Hourly Demand Forecast PerformanceNational Energy System Operator

    Source of the published day-ahead forecast and later corrected demand outturn, retained at the 12 August 2026 access point.

  3. NESO Open Data Licence v1.0National Energy System Operator

    Licence governing the retained source data. National Energy System Operator does not endorse this independent analysis.

What this answer does not establish

This article uses a Quanta Meridian example built without client data. Its calculation, process or threshold is not automatically the right rule for another organisation.

Return to all Insights