Skip to main content
Quanta Meridian logo

All examplesExample 13 of 15

Research ProjectsExample 3 of 3 in this collection

National demand forecasting · Python + R

Day-ahead Electricity Demand Forecast Review

A forecast analyst has tomorrow's half-hour demand curve and two proposed corrections. The question is not which model looks most sophisticated. It is whether either one improves the published forecast consistently enough to deserve controlled review.

Result after six rolling testsThe R correction earns further review. The Python correction does not clear the test.

R reduces mean absolute error by 2.1% and wins five of six origins. Python improves it by 0.7%, while its nominal 80% range covers 72.6% of outcomes.

For 2026-04-09/SP30, the later outturn is 25,045 MW. The published forecast remains the operating reference while R moves only to further controlled review.

This is a retrospective comparison. It does not establish future accuracy or authorise an operating change.

9 April 2026 · 48 half-hour periodsOne curve, two proposed corrections and the later outturn.
Day-ahead electricity demand forecast comparisonThe published forecast, R and Python corrections are compared with the later corrected demand outturn. Settlement period 30 is marked for detailed review.SP30 · review00:0012:0024:00Mobile day-ahead demand comparisonAll four series remain visible and are directly labelled. Settlement period 30 marks the largest published-forecast error in the final test.SP30ActualPublishedRPython00:0012:0024:00
Emphasise a forecast seriesPython 80% range
Public source rows checked
94,050
Held-out forecasts
8,066
Rolling origins
6
Published forecast MAE
697.9 MW

The working decision

Should the review process change?

The forecast analyst prepares the comparison for a demand-planning lead. Advancing the R correction means investigating it beside the current forecast under controlled conditions. It does not mean replacing the published curve or changing an operating instruction.

Weather, embedded generation, recurring calendar activity, special events and current system conditions remain outside this public model. Those are reasons for expert review, not details to conceal behind an aggregate score.

  1. 01Published forecast
  2. 02R and Python corrections
  3. 03Six seasonal tests
  4. 04Controlled review
  5. 05Demand-planning decision

Rolling-origin evaluation

Averages do not hide where a correction loses.

The tests begin in January, April, July and October 2025, then January and April 2026. Each model learns from the previous two years and predicts the next 28 days.

Mean absolute error across all six rolling forecast origins

Lower is better. Every origin covers the next 28 days, uses the same three forecasts and remains visible.

Pre-agreed review test

Accuracy, consistency and uncertainty must agree.

A proposed correction must improve MAE by at least 2%, win four origins, control bias and peak error, and place 75% to 85% of outcomes inside its nominal 80% range.

  1. 01
    Timing boundary

    Each model fits the previous 730 days and scores the following 28 days.

    Six chronological origins
  2. 02
    Leakage check

    The scoring rows contain only information available when the forecast is issued.

    Later outturn withheld
  3. 03
    Quarantined rows

    The ambiguous 30 October 2022 clock-change day is excluded before modelling.

    48 rows
  4. 04
    Python and R row agreement

    Both challengers use the same fold keys, published forecast and later outturn.

    8,066 predictions each
  5. 05
    Error and uncertainty

    MAE, bias, peak error and empirical interval coverage enter the selection test.

    Point and interval checks
R

R linear correction

Advance to controlled review
MAE change
2.1%
Origin wins
5 / 6
Interval coverage
79.4%
Checks passed
5 / 5
PY

Python boosted correction

Do not advance
MAE change
0.7%
Origin wins
4 / 6
Interval coverage
72.6%
Checks passed
3 / 5

One settlement-period trace

Follow the largest published-forecast miss in the final test.

2026-04-09/SP30 is deliberately an error review, not a typical period. The published forecast was issued on 8 April; the corrected outturn became available only after settlement period 30 on 9 April.

  1. 01
    Published forecast

    Available at the retained issue time

    20,708 MW
  2. 02
    R linear correction

    Compared, not adopted automatically

    20,580 MW
  3. 03
    Python boosted correction

    Compared, not adopted automatically

    20,537 MW
  4. 04
    Corrected demand outturn

    Observed after the target period

    25,045 MW
  5. 05
    Model-selection decision

    Advance the R linear correction to controlled review.

    r linear correction
  6. 06
    Demand-planning boundary

    Check weather, embedded generation, events and operating conditions before changing the process

    Human review required

Inspect the executed analysis

The page is a reading layer over retained Python, R and forecast rows.

The native report, R session record and executed Jupyter notebook reproduce the figures shown above. The source file remains checksum-locked at the 12 August 2026 access point.

Complete generated report · 8 checks passed · R 4.6.1 · Python 3.12

The full 1,440 × 3,725 report retains the curve, every origin, the gate result, settlement-period review and limitations in one inspectable image.

Open the complete forecast review report
Technical detailsView the source, feature contract, model comparison and limits

Source and grain

One NESO date and settlement period forms the business key. Forty-eight rows from the ambiguous 30 October 2022 clock-change day are quarantined in full; valid 46- and 50-period days remain.

R benchmark

Base R fits a linear correction to published-forecast residuals at each origin. Its model-based 80% prediction interval reaches 79.4% empirical coverage overall.

Python comparison

Scikit-learn gradient boosting predicts the residual median and 10th and 90th percentiles with fixed hyperparameters and seed 417. It does not use later outturn values at scoring time.

Decision limit

Further review must add operational inputs and governance. The example does not reproduce NESO's own method or direct system operation. It does not run as a live forecast.

Source and attribution

Supported by National Energy SO Open Data.

The retained resource is the Day Ahead Half Hourly Demand Forecast Performance file accessed 12 August 2026 under the NESO Open Data Licence v1.0. NESO does not endorse this independent analysis.

Related Research Brief

Why does a lower average error lead to more testing rather than replacement?

The brief separates the six-origin result, the complete selection gate and the operating decision so that “better” is not reduced to one average.

Read When is a better forecast strong enough to test further?

Discuss a forecasting review

Start with the issue time, the benchmark and the cost of a wrong decision.

A useful first discussion identifies what will be forecast, when the decision is made, what information exists then and how a proposed model must earn use.

Discuss a forecasting review