---
id: VOLT-HOME-WP-052
title: "How quickly does forecast skill decay from day one to day seven?"
slug: how-quickly-does-forecast-skill-decay-from-day-one-to-day-seven
description: "An exact-horizon, serving-lineage-matched measurement of how day-ahead price forecast MAE changes between the shortest and longest served horizons."
published: 2026-08-30
cluster: "Forecast verification and uncertainty"
status: measured
evidence: /research-data/home-papers/how-quickly-does-forecast-skill-decay-from-day-one-to-day-seven.json
figure: /research-media/home-papers/how-quickly-does-forecast-skill-decay-from-day-one-to-day-seven.webp
figure_alt: "Chart for How quickly does forecast skill decay from day one to day seven?: MAE increase from shortest to longest exact horizon, shown as d1, d2, d3, d4, d5, d6, d7."
source_ids:
  - volt-promotion-policy
  - volt-v61-card
  - volt-research-content
  - google-article
  - entsoe-sdac
peer_reviewed: false
---

## Abstract

Seven-day forecast claims can be misleading when short and long horizons are mixed. This paper measures the difference between the mean absolute error (MAE) at the longest and shortest exact horizons in the frozen serving record. The measured increase is 25.375216403577852 EUR/MWh, based on 17750 lineage-matched accuracy rows. The figure preserves the complete d1-through-d7 sequence: 60.718384096109844, 75.86942822966508, 89.23258108192621, 90.66425893915756, 89.97785701917131, 89.86564108091223, and 86.0936004996877 EUR/MWh.

Eligibility is strict: a row counts as live-service evidence only when its model version exactly matches the immutable serving-lineage record for the same bidding zone, delivery date, and horizon bucket. Development models, reconstructed proxies, and versions promoted for a different objective are excluded. The result says that the longest served horizon was less accurate on average in absolute-error terms than the shortest served horizon. It does not show that error rises smoothly every day, nor that every zone has the same decay curve. It is descriptive, model-assisted, and not peer reviewed.

## Plain-language answer

In this record, moving from the shortest exact served horizon to the longest exact served horizon adds 25.375216403577852 EUR/MWh to mean absolute error. That is the direct answer supported by the evidence. A day-one score should not be presented as evidence of day-seven skill, because the information available to the model and the uncertainty around the target change with lead time.

The result is a fleet-level endpoint contrast. It does not mean that each successive day adds an equal amount of error. Some zones may have flatter profiles; some dates may become difficult abruptly; and horizon composition can differ across markets. Users planning a device several days ahead should therefore inspect the score for the exact horizon they intend to use. A good next-day curve is not a warranty for a week-ahead curve.

## Research question

The preregistered question asks how quickly forecast skill decays from day one to day seven. The operational estimand is the longest-horizon mean MAE minus the shortest-horizon mean MAE among rows that are both scored and serving-lineage matched. The word “skill” is used cautiously here: the headline is an error difference, not a normalized skill score relative to a common baseline.

The exact-horizon requirement matters because a seven-day API can contain materially different information sets. The day-one forecast can use more recent prices, weather runs, and system observations than the day-seven forecast. ENTSO-E's [Single Day-ahead Coupling (SDAC)](https://www.entsoe.eu/network_codes/cacm/implementation/sdac/) explains the coupled institutional setting for the target prices, but it does not imply identical predictability across horizons or zones. This paper measures the served record rather than deriving predictability from the market design.

## Data and provenance

The evidence is the frozen public JSON named in the frontmatter. Its cutoff is `2026-08-30T00:00:00Z`. The family declares `forecast_accuracy`, `forecast_serving_lineage`, `risk_accuracy`, `generation_mix`, and `zone_temp_weighted`; the endpoint calculation uses the first two. Eligible rows have non-null MAE and an exact match between the scored `model_version` and the `served_model_version` for zone, delivery date, and horizon bucket.

The source snapshot covers the declared production windows without exposing the temporary aggregate input. Its SHA-256 is `7e97489fc8528c8cc8c38830e05b48d949ce1f67b98575dff26f5d7c321e4c67`. The analysis code, protocol, registry, and source-registry hashes are `57c57de79cdab2b5b6d6c54c485cb5162598c5ba0b0bfe995da40d75e6c52ba9`, `adb36bf6b447af9f96339249b8becaefc20422499cca1977242866347a97bd4b`, `7bcb91d7476d0a69fe9fa75a5c7782f8117e0153f82f9112b7e1d307d3943717`, and `07949550ac443ff673fda5c0209b99f137544f3ffecf6775f109bb9d09663bd6`. The evidence and figure hashes are `c0cf5879c6b637b66b38131368f9a9f2ce4dcdf31817120d763731e3aee4ab6a` and `ccd3a051723f3153b11118d5d363d325cd9332bba8c215eab840f01ebcf52709`. Extraction was SELECT-only in a read-only transaction with a 180-second statement timeout.

## Method

The analysis groups all eligible MAE rows by integer `horizon_days`. It calculates the arithmetic mean MAE within each nonzero exact horizon, sorts those horizon means, and subtracts the first from the last. The chart labels d1, d2, d3, d4, d5, d6, and d7, establishing that the displayed profile covers the full served sequence. The headline reports only the endpoint increase.

This is deliberately different from choosing the best model at each horizon after seeing outcomes. The [Objective-specific seven-day forecast promotion](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/OBJECTIVE-PROMOTION-POLICY.md) requires exact-horizon evidence and keeps point, quantile, external-challenger, and customer-production objectives separate. The [Volt 6.1 cross-zone price model card](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/V6.1-MODEL-CARD.md) describes a development programme spanning exact horizons, but that version failed its registered gate and never became serving evidence. Its offline proxy scores are not inserted into this live-service curve.

## Results

The shortest-to-longest exact-horizon MAE increase is 25.375216403577852 EUR/MWh. The sample field reports 17750 underlying lineage-matched MAE rows. Positive endpoint change means the longest exact horizon had higher average absolute error than the shortest exact horizon in the frozen serving record.

The regenerated evidence publishes no uncertainty interval for the 25.375216403577852 endpoint contrast and identifies the interval method as “not reported for this estimand.” No significance claim is made.

The figure contains the horizon means used for the visual profile: d1 60.718384096109844, d2 75.86942822966508, d3 89.23258108192621, d4 90.66425893915756, d5 89.97785701917131, d6 89.86564108091223, and d7 86.0936004996877 EUR/MWh. These values show that the profile is non-monotonic even though the endpoint difference is positive.

## Robustness and placebo checks

Serving-lineage matching is the principal anti-survivorship control. A scored version that was not the recorded serving choice for its zone, date, and horizon cannot improve or worsen the live curve. Exact-horizon grouping is the second control: it prevents a dense collection of short-horizon rows from being described as seven-day evidence. Objective separation prevents a point-only benchmark model from being silently treated as a probabilistic customer forecast.

Several useful placebos are not in the frozen evidence. There is no shuffled-horizon test, balanced zone-date panel, same-model-only profile, or comparison with a horizon-invariant persistence baseline. No uncertainty interval is reported, so serial dependence and cross-horizon dependence are not quantified. The multiplicity statement correctly limits this output to a descriptive result without an unadjusted inferential claim.

## Limitations

MAE is denominated in EUR/MWh and can be influenced by zone price scale and event regimes. Pooling zones means a change in the mix of markets observed at longer horizons could contribute to the endpoint difference. The public evidence does not provide matched counts by zone and horizon, so a balanced within-zone decay estimate cannot be reconstructed from the paper-level aggregate.

The headline does not reveal shape. Error could rise steadily, remain flat before a late jump, or vary non-monotonically. Nor does the analysis separate forecast versions. Serving lineage makes each row historically honest, but the serving model can change over time. The endpoint contrast therefore measures the production service as operated, not the intrinsic decay curve of one fixed algorithm.

Finally, “skill decay” is broader than MAE growth. Proper probabilistic scores, interval coverage, calibration, and decision utility can decay differently. A model can retain median accuracy while its tails become underdispersed. The objective-specific policy requires those outputs to be evaluated separately; this paper does not collapse them into one score.

The endpoint statistic also cannot establish when additional information becomes valuable. Forecast updates between horizons may arrive on different schedules, and a horizon label alone does not identify the weather or market vintages available to the model. A causal account of decay would need those receipt clocks as covariates.

## Practical implication

Household automation should bind each decision to the relevant exact horizon. A controller scheduling tomorrow can use the day-one record; a planner reserving energy several days ahead should use the matching longer-horizon score and wider uncertainty. The measured 25.375216403577852 EUR/MWh endpoint increase is a warning against applying the shortest-horizon accuracy badge to the entire week.

For model operators, the result supports per-horizon promotion and rollback. A candidate that improves d1 but degrades d7 should not replace a seven-day customer contract. This is why the production-main policy requires the complete horizon set and why point-only external evidence remains a separate objective. The [Voltcast Research Content Plan](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/RESEARCH-CONTENT-PLAN.md) also requires the data window and losing results to remain visible instead of compressing them into a favorable headline.

## Reproducibility

A reproducer should verify the evidence file and manifest hash, then verify the snapshot, code, protocol, and registry SHA-256 values. Join scored accuracy to serving lineage on zone, delivery date, horizon bucket, and exact model version. Reject non-matches and null MAEs. Group by nonzero `horizon_days`, calculate each mean, sort horizons, and subtract the shortest mean from the longest.

To strengthen—not restate—the result, a new protocol could restrict to a common set of zone-dates with all seven horizons, estimate paired within-zone changes, and block-resample delivery weeks. It could also compare fixed model versions and common persistence baselines. Those would answer different estimands and must be published as new evidence. Google Search Central's assigned [Article structured data](https://developers.google.com/search/docs/appearance/structured-data/article) source supports exposing the stable title, date, image, and evidence linkage in the public rendering.

## Disclosure

Analysis and drafting were model-assisted; sources, code, assumptions, and evidence hashes are disclosed. The model wrote explanatory prose around the frozen aggregate and did not supply empirical values. This working paper is not peer reviewed.

Volt has zero live traders and zero live capital. C0R is the only paper strategy. Production forecasting is a non-trading service. This horizon analysis is not a trading system and does not authorize capital. It is not trading advice or financial advice, and it does not estimate a household's retail bill or promise savings.

## References

1. Voltcast. “Objective-specific seven-day forecast promotion.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/OBJECTIVE-PROMOTION-POLICY.md
2. Voltcast. “Volt 6.1 cross-zone price model card.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/V6.1-MODEL-CARD.md
3. Voltcast. “Voltcast Research Content Plan.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/RESEARCH-CONTENT-PLAN.md
4. Google Search Central. “Article structured data.” https://developers.google.com/search/docs/appearance/structured-data/article
5. ENTSO-E. “Single Day-ahead Coupling (SDAC).” https://www.entsoe.eu/network_codes/cacm/implementation/sdac/
