---
id: VOLT-HOME-WP-051
title: "What makes one European bidding zone harder to forecast than another?"
slug: what-makes-one-european-bidding-zone-harder-to-forecast-than-another
description: "A serving-lineage-matched comparison of day-ahead price forecast error across European bidding zones, with fleet averages kept separate from zone-level evidence."
published: 2026-08-30
cluster: "Forecast verification and uncertainty"
status: measured
evidence: /research-data/home-papers/what-makes-one-european-bidding-zone-harder-to-forecast-than-another.json
figure: /research-media/home-papers/what-makes-one-european-bidding-zone-harder-to-forecast-than-another.webp
figure_alt: "Chart for What makes one European bidding zone harder to forecast than another?: cross-zone dispersion in lineage-matched MAE, shown as UA, LT, SE4, LV, EE."
source_ids:
  - volt-promotion-policy
  - volt-v61-card
  - volt-research-content
  - google-article
  - entsoe-sdac
peer_reviewed: false
---

## Abstract

Forecast difficulty is not uniform across European bidding zones. In the frozen public evidence for this paper, the population standard deviation of zone-level mean absolute error (MAE) is 380.676952413692 EUR/MWh across 46 zones. The calculation admits only forecast-accuracy rows whose model version exactly matches the immutable serving-lineage decision for the same zone, delivery date, and horizon bucket. It therefore describes models that were actually selected for service, rather than whichever historical model looks best after outcomes are known.

The result establishes dispersion, not its cause. Price scale, event frequency, market structure, data completeness, model age, horizon mix, and the objectives under which a version was promoted can all differ by zone. The evidence does not decompose those explanations. It also does not authorize a fleet average to stand in for any local market. The chart identifies UA, LT, SE4, LV, and EE as the five displayed high-MAE labels, with values of 2641.2191664, 63.64336851851852, 60.15685498575499, 58.491252488687785, and 56.583720138888886 EUR/MWh. This is a model-assisted, non-peer-reviewed working paper about a non-trading production forecasting service.

## Plain-language answer

One European bidding zone can be harder to forecast than another because the local price process and the evidence available to a serving model are not interchangeable. The measured answer is that the differences are large enough that a single Europe-wide accuracy number is not an adequate description: the cross-zone dispersion in lineage-matched MAE is 380.676952413692 EUR/MWh in the frozen sample.

That number should not be read as a permanent ranking of countries, a measure of household bills, or proof that any named market is structurally inefficient. It is a dispersion statistic over the serving record available before the publication cutoff. A zone can look difficult because a small number of very large errors raise its average, because the served horizon mix is longer, because its currently eligible model is older, or because the local price scale is wider. Conversely, an apparently easy zone can have fewer difficult regimes represented in the matched sample. Local consumers should therefore inspect their own zone and horizon instead of relying on the fleet mean.

## Research question

The preregistered question is: what makes one European bidding zone harder to forecast than another? The measurable part is whether lineage-matched forecast error differs across zones. The stronger causal wording—what *makes* the difference—would require a separate design that holds horizon, date, model objective, price scale, and regime exposure constant while varying candidate drivers.

This paper answers only the first part. It estimates each zone's mean MAE from eligible rows and summarizes the dispersion of those zone means. It does not attribute the dispersion to market coupling, renewable penetration, temperature, liquidity, or model architecture. ENTSO-E's [Single Day-ahead Coupling (SDAC)](https://www.entsoe.eu/network_codes/cacm/implementation/sdac/) supplies the institutional context that these are coupled cross-zonal day-ahead markets, but coupling does not make their realized price distributions or local constraints identical. Association is not mechanism, and a ranking is not a causal model.

## Data and provenance

The paper is bound to the frozen JSON at the evidence URL in the frontmatter. Its publication cutoff is `2026-08-30T00:00:00Z`. The declared source contracts are `forecast_accuracy`, `forecast_serving_lineage`, `risk_accuracy`, `generation_mix`, and `zone_temp_weighted`. This particular outcome uses the first two: scored forecast rows and the immutable record of which version was served. Only rows with non-null MAE and an exact lineage match enter the result.

The aggregate snapshot was extracted in a read-only transaction with a 180-second statement timeout. Its SHA-256 is `7e97489fc8528c8cc8c38830e05b48d949ce1f67b98575dff26f5d7c321e4c67`; the analysis-code hash is `57c57de79cdab2b5b6d6c54c485cb5162598c5ba0b0bfe995da40d75e6c52ba9`; the protocol hash is `adb36bf6b447af9f96339249b8becaefc20422499cca1977242866347a97bd4b`; the registry hash is `7bcb91d7476d0a69fe9fa75a5c7782f8117e0153f82f9112b7e1d307d3943717`; and the source-registry hash is `07949550ac443ff673fda5c0209b99f137544f3ffecf6775f109bb9d09663bd6`. The evidence and figure hashes are `cc21655b8982dc5444e6c1811d84b53744009fcc8f753cc8551f4ef031ece40d` and `8937efc16e0ffc78b9fc72d570a14213ae1c767bd94173ea0719ca4dd1b0f92c`. These identifiers make the public aggregate auditable without publishing row-level production data.

## Method

For every lineage-matched accuracy row, the analysis groups MAE by bidding-zone code. It computes the arithmetic mean within each zone. It then calculates the population standard deviation of the resulting zone means. This gives the headline cross-zone dispersion. The sample size of 46 is the number of zone-level values, not the number of delivery intervals or customers.

The method intentionally avoids selecting one globally best historical model. Voltcast's [Objective-specific seven-day forecast promotion](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/OBJECTIVE-PROMOTION-POLICY.md) separates point, quantile, external-challenger, and customer-API policies. A point-only candidate cannot silently replace a probabilistic customer model, and a benchmark submission is not a serving decision. The [Volt 6.1 cross-zone price model card](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/V6.1-MODEL-CARD.md) is especially relevant: Volt 6.1 is development evidence, failed its registered replacement gate, and is not serving. Its model families and objectives must not be blended into this live-service statistic merely because they exist in the repository.

## Results

Across 46 zone-level mean-MAE observations, the measured cross-zone population standard deviation is 380.676952413692 EUR/MWh. The frozen interpretation is that zone difficulty differs materially, so one fleet average cannot represent every household market. The displayed high-MAE labels and values are UA at 2641.2191664, LT at 63.64336851851852, SE4 at 60.15685498575499, LV at 58.491252488687785, and EE at 56.583720138888886 EUR/MWh.

The regenerated evidence records `bootstrap_95_interval` as null and the interval method as “not reported for this estimand.” No uncertainty interval accompanies the population-standard-deviation headline. The result is descriptive and carries no unadjusted significance claim.

## Robustness and placebo checks

The strongest robustness control is serving lineage. Every included error must match the version recorded as served for the same zone, date, and horizon bucket. That blocks a common placebo victory in which an analyst tests many archived models and retrospectively assigns each outcome to the winner. Objective separation is another control: point, probabilistic, external-benchmark, and production-main evidence are not pooled as if they answered the same question.

The family protocol calls for blocked uncertainty intervals for inferential claims, but this descriptive estimand reports no interval. Serial dependence is therefore not quantified. No causal placebo—such as shuffled zone labels, matched price-scale strata, or held-out regime windows—is reported in the evidence. Those checks would be appropriate for a follow-up, but claiming they were run would be false. Within-family Holm control is reserved for inferential claims; this paper makes none.

## Limitations

MAE is scale-dependent. A high-price or spike-prone zone can produce a larger EUR/MWh error even if its relative accuracy is acceptable. The evidence does not normalize by price level, interquartile range, or a common persistence baseline. It also aggregates all eligible exact horizons represented in each zone. Unequal horizon composition can therefore contribute to the observed dispersion.

The 46 zone means may have unequal numbers of matched days. The evidence does not publish per-zone counts, individual MAEs, or a balanced common-date panel. It cannot distinguish an inherently difficult market from a temporary regime or an older served model. The five displayed codes are not a durable league table. Most importantly, this design measures heterogeneity but does not identify why it exists. Generation and temperature contracts are declared for the family, yet this result does not use them in a driver decomposition.

## Practical implication

Forecast users should evaluate the local zone, exact horizon, objective, and served model version that match their decision. A Europe-wide badge can be useful as an operational summary, but it should never replace the local scorecard. For household automation, the relevant question is whether tomorrow's curve for the household's own bidding zone is sufficiently complete, timely, and calibrated for the intended device—not whether the fleet average improved.

Model operators should retain per-zone gates and rollback identities. A strong result in one zone should not authorize promotion in another, and a strong point forecast should not authorize a quantile-serving change. This is consistent with the objective-specific promotion policy and with the [Voltcast Research Content Plan](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/RESEARCH-CONTENT-PLAN.md), which requires losses and misses to be published with the same prominence as wins.

## Reproducibility

Reproduction starts from the registered paper identity, the frozen evidence JSON, and its provenance. A reproducer should verify the paper ID, slug, evidence and figure hashes from the manifest, snapshot hash, protocol hash, registry hash, source-registry hash, and analysis-code hash before recalculating anything. It should then join scored accuracy to serving lineage on zone, delivery date, horizon bucket, and exact served model version; reject unmatched rows; calculate mean MAE per zone; and take the population standard deviation across those means.

A stronger rerun should add a common-date and common-horizon panel, publish zone counts, and use a genuine blocked bootstrap around the dispersion statistic. Those would be new outputs and must not be backfilled into this frozen result. The public page's article metadata follows the assigned Google Search Central source, [Article structured data](https://developers.google.com/search/docs/appearance/structured-data/article), so title, publication date, image, and evidence identity remain machine-readable as well as visible.

## Disclosure

Analysis and drafting were model-assisted. The language model organized and explained the frozen result; it was not the source of any empirical number. Sources, method, assumptions, limitations, and evidence hashes are disclosed. This working paper is not peer reviewed.

Volt has zero live traders and zero live capital. C0R is the only paper strategy. Voltcast production forecasting is a non-trading service, and this cross-zone accuracy analysis neither opens nor evaluates a trading strategy. This paper is not trading advice, financial advice, a retail-bill forecast, or evidence from customer bills.

## References

1. Voltcast. “Objective-specific seven-day forecast promotion.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/OBJECTIVE-PROMOTION-POLICY.md
2. Voltcast. “Volt 6.1 cross-zone price model card.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/V6.1-MODEL-CARD.md
3. Voltcast. “Voltcast Research Content Plan.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/RESEARCH-CONTENT-PLAN.md
4. Google Search Central. “Article structured data.” https://developers.google.com/search/docs/appearance/structured-data/article
5. ENTSO-E. “Single Day-ahead Coupling (SDAC).” https://www.entsoe.eu/network_codes/cacm/implementation/sdac/
