---
id: VOLT-HOME-WP-054
title: "What happens to probabilistic coverage during price spikes?"
slug: what-happens-to-probabilistic-coverage-during-price-spikes
description: "A lineage-matched audit showing how nominal 95% coverage changes in the highest-MAE forecast cells, with a clear warning that the frozen regime is forecast error rather than realized price spikes."
published: 2026-08-30
cluster: "Forecast verification and uncertainty"
status: measured
evidence: /research-data/home-papers/what-happens-to-probabilistic-coverage-during-price-spikes.json
figure: /research-media/home-papers/what-happens-to-probabilistic-coverage-during-price-spikes.webp
figure_alt: "Chart for What happens to probabilistic coverage during price spikes?: coverage95 change in the MAE spike decile, shown as normal cells, spike decile."
source_ids:
  - volt-promotion-policy
  - volt-v61-card
  - volt-research-content
  - google-article
  - entsoe-sdac
peer_reviewed: false
---

## Abstract

Probabilistic coverage falls sharply in the most error-prone part of the lineage-matched serving record. The frozen result defines its “spike” regime as the top decile of MAE, not as a top decile of realized electricity prices. Within that implemented regime, `coverage_95` changes by -0.3043698197662636 relative to the remaining cells. The spike subset contains 1310 coverage observations.

This is evidence of conditional undercoverage when point error is already high. It is not direct evidence about price spikes, because the grouping variable is forecast MAE. A high-MAE cell can arise at an ordinary price level, and a large realized price can still be forecast accurately. The distinction prevents a circular result from being marketed as an event study. The paper therefore reports the measured high-error-regime finding and treats the title's realized-price-spike question as unresolved by this frozen design. Only exact serving-lineage matches count as live-service evidence; development models and objective-mismatched versions do not.

## Plain-language answer

When the serving system has its largest absolute errors, its reported wide interval covers outcomes much less often. The measured change in `coverage_95` is -0.3043698197662636 for the MAE spike decile versus other cells. That is a large deterioration in the coverage fraction.

However, the evidence did not select hours or days because electricity prices themselves spiked. It selected forecast cells because their errors were in the highest decile. The result therefore answers “what happens to coverage when forecasts miss badly?” Coverage deteriorates. It does not cleanly answer “what happens when realized prices are unusually high?” A genuine price-spike study must define the event using realized prices before examining forecast error or coverage.

## Research question

The intended question concerns whether probabilistic bands remain reliable during extreme price events. Such a study would ordinarily define an event threshold from the realized price distribution, freeze that definition, and compare interval inclusion inside and outside the event set. It might distinguish positive spikes, negative prices, ramps, and sustained high-price episodes because their forecast mechanisms differ.

The implemented estimand is narrower and different. It calculates a threshold at the 0.9 quantile of lineage-matched MAE, calls rows at or above that threshold the spike subset, and compares their mean `coverage_95` with all rows below the threshold. This conditions on the error whose uncertainty performance is being assessed. It is useful as a stress diagnostic, but not an independent price-event classification. ENTSO-E's [Single Day-ahead Coupling (SDAC)](https://www.entsoe.eu/network_codes/cacm/implementation/sdac/) provides the coupled day-ahead market context; it does not supply the missing realized-price event definition.

## Data and provenance

The public JSON named in the frontmatter is the evidence of record. Its publication cutoff is `2026-08-30T00:00:00Z`. The family contracts are `forecast_accuracy`, `forecast_serving_lineage`, `risk_accuracy`, `generation_mix`, and `zone_temp_weighted`. The implemented outcome uses non-null MAE and coverage-95 fields from accuracy rows whose model version exactly matches immutable serving lineage for zone, delivery date, and horizon bucket.

The production aggregate was extracted in a SELECT-only, read-only transaction with a 180-second statement timeout. The snapshot hash is `7e97489fc8528c8cc8c38830e05b48d949ce1f67b98575dff26f5d7c321e4c67`. The analysis-code, protocol, registry, and source-registry hashes are `57c57de79cdab2b5b6d6c54c485cb5162598c5ba0b0bfe995da40d75e6c52ba9`, `adb36bf6b447af9f96339249b8becaefc20422499cca1977242866347a97bd4b`, `7bcb91d7476d0a69fe9fa75a5c7782f8117e0153f82f9112b7e1d307d3943717`, and `07949550ac443ff673fda5c0209b99f137544f3ffecf6775f109bb9d09663bd6`. The evidence and figure hashes are `2ffbe2facc0cdb991fa57e3ce91afb2480e78830e988bb3b13ec398f0c0da005` and `02981dc81a69a55b45261a0029ceeea29d2f1be6fddf1210937a1dd86ad38592`. The temporary snapshot is not a publication artifact; only aggregate evidence is public.

## Method

The analysis first forms the eligible serving sample. It takes the 0.9 quantile of MAE over that sample. Rows at or above the resulting threshold enter the high-error subset, provided `coverage_95` is available; rows below it enter the comparison subset. The primary statistic is mean coverage in the high-error subset minus mean coverage in the comparison subset. A negative value means coverage is worse among high-error rows.

The method is version-honest but not model-homogeneous. It can include multiple serving versions over time, provided each row was the actual serving choice. The [Objective-specific seven-day forecast promotion](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/OBJECTIVE-PROMOTION-POLICY.md) prevents a point-only candidate from standing in for a probabilistic serving model. The [Volt 6.1 cross-zone price model card](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/V6.1-MODEL-CARD.md) describes a failed development programme that never served; none of its proxy outcomes enter this live-service result.

## Results

The measured `coverage_95` change in the MAE spike decile is -0.3043698197662636. The high-error subset contributes 1310 observations. The figure reports 77.54691327433628% coverage in normal cells and 47.10993129770992% in the spike decile; their difference is -30.43698197662636 percentage points, equivalent to the primary fraction. The sign means the nominal wide interval includes realized outcomes less often among rows with the largest point errors than among the remaining eligible rows.

The regenerated evidence records `bootstrap_95_interval` as null and the interval method as “not reported for this estimand.” No uncertainty interval accompanies the -0.3043698197662636 change, so it cannot be used to claim significance for the difference.

No realized-price threshold, price-event count, or positive-spike coverage statistic appears in the JSON. The paper does not infer one from the title, figure styling, or other families. The empirical conclusion is limited to high-MAE conditional coverage.

## Robustness and placebo checks

Serving-lineage matching is the principal robustness control. It blocks retrospective selection of whichever archived model produced the best bands on known outcomes. Exact matching also maintains the horizon bucket and model identity associated with the serving decision. Objective separation prevents a sharp point model without valid uncertainty outputs from being counted as a probabilistic champion.

The most important missing placebo is an independently defined price-spike regime. A suitable follow-up could freeze realized-price quantiles on a training period, apply them to a held-out period, and compare with matched non-event days. Other useful checks include equal-size random-day subsets, positive versus negative extremes, zone-fixed effects, and a common set of model versions. None is in this evidence.

The family protocol calls for blocked uncertainty intervals for inferential claims, but this descriptive output reports none. Serial dependence across dates and related horizons is therefore not quantified. Within-family Holm control is mentioned for inferential claims, yet this output properly makes no unadjusted significance claim.

## Limitations

Conditioning on MAE creates a partly mechanical relationship. When realized outcomes fall far from point forecasts, finite forecast intervals may also miss more often. The result quantifies that co-occurrence in the serving record but does not explain whether intervals failed to widen, whether the median shifted, or whether a particular tail was responsible.

The coverage field can be aggregated at a cell level with differing numbers of periods. The public evidence does not expose weights, zone-horizon strata, or per-version counts. A pooled mean can overrepresent cells that are numerous rather than economically or temporally comparable. It also does not report interval width, so calibration and sharpness cannot be evaluated jointly.

Most importantly, MAE spikes are not price spikes. This semantic mismatch prevents the title's full question from receiving a causal or event-specific answer. The result should be used as a reliability warning and as a design critique, not relabeled as realized-price evidence.

Selection on high MAE also removes any clean temporal ordering between regime and failure. The subset is known only after the realized price is available and error has been scored. It cannot be used prospectively as an alert without a separate, issue-time predictor of difficult conditions. A useful operational study would define observable pre-issue covariates, freeze a difficulty classifier, and then test whether conditional coverage deteriorates in its flagged rows. That is a different design from sorting realized misses and must not borrow this paper's result as prospective proof.

## Practical implication

Automation should become more conservative when forecast conditions resemble historically difficult regimes. A wide nominal band that under-covers precisely where point error is high cannot be treated as a device-safety boundary. Controllers should retain hard physical limits, freshness checks, fallback schedules, and user-defined risk tolerances independent of the forecast interval.

For model governance, this result supports cell-level coverage gates and rollback. A model should not promote because its average MAE improves if its probabilistic coverage collapses in difficult cases. That principle is explicit in the objective-specific promotion policy. The [Voltcast Research Content Plan](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/RESEARCH-CONTENT-PLAN.md) further requires misses to be published with the same prominence as wins; the -0.3043698197662636 result is therefore not hidden behind an average score.

## Reproducibility

Verify the evidence and manifest identities, then the snapshot, code, protocol, and registry hashes. Join scored accuracy to serving lineage by zone, delivery date, horizon bucket, and exact model version. Retain non-null MAE rows. Compute the 0.9 MAE quantile over that sample, separate rows at or above it from rows below it, retain non-null coverage-95, and subtract the comparison mean from the high-error mean.

A correct realized-price-spike extension must be preregistered separately. It should define price events from a clock-safe realized-price series, avoid choosing thresholds on the evaluation outcomes when possible, preserve zone and season comparability, and use blocked uncertainty. It must not overwrite this high-error result. The public rendering uses Google Search Central's assigned [Article structured data](https://developers.google.com/search/docs/appearance/structured-data/article) source for discoverable title, image, date, and evidence metadata.

## Disclosure

Analysis and drafting were model-assisted. The model helped explain the frozen statistic and surfaced the mismatch between the title and implemented event definition. It did not create empirical observations. Sources, code, assumptions, limitations, and hashes are disclosed. This working paper is not peer reviewed.

Volt has zero live traders and zero live capital. C0R is the only paper strategy. Production forecasting is a non-trading service. This paper is not trading advice, financial advice, a customer-savings estimate, or authorization to promote a model.

## References

1. Voltcast. “Objective-specific seven-day forecast promotion.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/OBJECTIVE-PROMOTION-POLICY.md
2. Voltcast. “Volt 6.1 cross-zone price model card.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/V6.1-MODEL-CARD.md
3. Voltcast. “Voltcast Research Content Plan.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/RESEARCH-CONTENT-PLAN.md
4. Google Search Central. “Article structured data.” https://developers.google.com/search/docs/appearance/structured-data/article
5. ENTSO-E. “Single Day-ahead Coupling (SDAC).” https://www.entsoe.eu/network_codes/cacm/implementation/sdac/
