---
id: VOLT-HOME-WP-060
title: "Why backtest accuracy and live accuracy diverge"
slug: why-backtest-accuracy-and-live-accuracy-diverge
description: "A serving-lineage account of live forecast evidence, showing a small MAE improvement over persistence while keeping retrospective backtests and prospective service as different estimands."
published: 2026-08-30
cluster: "Forecast verification and uncertainty"
status: measured
evidence: /research-data/home-papers/why-backtest-accuracy-and-live-accuracy-diverge.json
figure: /research-media/home-papers/why-backtest-accuracy-and-live-accuracy-diverge.webp
figure_alt: "Chart for Why backtest accuracy and live accuracy diverge: lineage-matched MAE improvement over persistence, shown as served model, persistence."
source_ids:
  - volt-promotion-policy
  - volt-v61-card
  - volt-research-content
  - google-article
  - entsoe-sdac
peer_reviewed: false
---

## Abstract

Backtests and live forecasts diverge because they do not necessarily use the same information, model-selection clock, version, objective, or operational constraints. The frozen empirical result in this paper measures only the live side: among 17750 serving-lineage-matched rows, persistence MAE minus served-model MAE is 0.7614007042253661 EUR/MWh. The positive value means the recorded serving model has slightly lower pooled MAE than the recorded persistence baseline.

The evidence does not contain a paired backtest-versus-live accuracy difference. It therefore cannot quantify the size of divergence named in the title. Instead, it demonstrates the discipline required before such a comparison: live evidence must match the exact model selected before outcomes, while reconstructed or development backtests remain separately labeled. Volt 6.1 illustrates the distinction because its historical comparator was a proxy and its registered replacement gate failed; it never became serving evidence. The paper answers the “why” from the governing contracts and reports the available live benchmark result without pretending that it is a direct backtest gap.

## Plain-language answer

A backtest asks how a frozen method would have performed on historical data under reconstructed information rules. A live score asks how the version actually selected and issued performed under real operations. Those can differ because historical data may be cleaner, publication times may be approximated, model choice may use hindsight, and production can encounter missing inputs or version transitions.

In the actual lineage-matched evidence here, the served model improves MAE over persistence by 0.7614007042253661 EUR/MWh across 17750 rows. That is a live-service benchmark comparison. It does not say that the backtest was better or worse by that amount. The safe conclusion is that live accuracy must be measured from immutable serving decisions, and any backtest-live comparison must match the same version, horizon, objective, and data clock.

## Research question

The title asks why retrospective accuracy and prospective operational accuracy can diverge. The corresponding empirical design would pair a frozen backtest prediction and an actually issued prediction for the same zone, target, horizon, and model definition, then decompose differences into source vintage, feature availability, artifact version, and operational completeness.

The implemented outcome is narrower. It compares MAE of serving-lineage-matched forecasts with their recorded persistence baseline. That is a valid live performance statistic, but it has no backtest arm. ENTSO-E's [Single Day-ahead Coupling (SDAC)](https://www.entsoe.eu/network_codes/cacm/implementation/sdac/) establishes the coupled day-ahead price target context. It does not solve the historical information problem: backtests still need issue-time source vintages and exact operational clocks to emulate what was knowable.

## Data and provenance

The frozen JSON at the frontmatter evidence URL has publication cutoff `2026-08-30T00:00:00Z`. The family contracts are `forecast_accuracy`, `forecast_serving_lineage`, `risk_accuracy`, `generation_mix`, and `zone_temp_weighted`. This outcome uses non-null MAE rows with exact serving-lineage matches and non-null `baseline_mae`.

The production aggregate was extracted through a SELECT-only, read-only transaction with a 180-second statement timeout. Snapshot SHA-256 is `7e97489fc8528c8cc8c38830e05b48d949ce1f67b98575dff26f5d7c321e4c67`. Analysis-code, protocol, registry, and source-registry hashes are `57c57de79cdab2b5b6d6c54c485cb5162598c5ba0b0bfe995da40d75e6c52ba9`, `adb36bf6b447af9f96339249b8becaefc20422499cca1977242866347a97bd4b`, `7bcb91d7476d0a69fe9fa75a5c7782f8117e0153f82f9112b7e1d307d3943717`, and `07949550ac443ff673fda5c0209b99f137544f3ffecf6775f109bb9d09663bd6`. The evidence and figure hashes are `d1d4de89e3607ea2e9b1b6ada9b0623d56f9b27b846e2a4197463329f6bfd741` and `68bbfe49ab0423c7a87bc8de461492e0546b5ed78b40954e7f6aadf5525b36d8`.

Serving lineage matches zone, delivery date, horizon bucket, and exact model version. This is the minimum identity needed to call an accuracy row live-service evidence in this paper.

## Method

The analysis first selects rows with non-null MAE that exactly match immutable serving lineage. It forms one list of served-model MAEs and another list of non-null persistence baseline MAEs. It calculates each list's arithmetic mean and subtracts served-model mean MAE from baseline mean MAE. Positive improvement indicates lower error for the served model.

This is not a matched row-by-row subtraction in the disclosed code; it is a difference between pooled means. The evidence does not publish whether every served MAE has a corresponding baseline value, so the precise denominator relationship should not be assumed beyond the frozen implementation.

The [Objective-specific seven-day forecast promotion](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/OBJECTIVE-PROMOTION-POLICY.md) explains why one accuracy concept cannot authorize every serving surface. Point, quantile, external-challenger, and production-main objectives have different output contracts. The [Volt 6.1 cross-zone price model card](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/V6.1-MODEL-CARD.md) explains why development evidence is not live evidence: its historical comparator was reconstructed rather than actual serving lineage, its registered gate failed, and it did not serve.

## Results

The measured lineage-matched MAE improvement over persistence is 0.7614007042253661 EUR/MWh. The figure reports served-model MAE of 87.14137855774648 EUR/MWh and persistence MAE of 87.90277926197184 EUR/MWh, whose difference reproduces the headline. The sample field reports 17750 served-model MAE observations. A positive value means the served forecasts have lower pooled mean MAE than persistence in the frozen record.

The regenerated evidence records `bootstrap_95_interval` as null and the interval method as “not reported for this estimand.” No uncertainty interval accompanies the 0.7614007042253661 improvement, and it is not used to claim significance.

No backtest MAE, backtest-live difference, version-specific gap, or source-vintage decomposition is in the JSON. The paper does not fabricate those results. Its empirical contribution is the lineage-matched live comparison with persistence.

## Robustness and placebo checks

Exact serving lineage is the main anti-survivorship control. It prevents a historical winner from being substituted for the version that users actually received. Keeping persistence as a recorded baseline avoids comparing the served model with an undefined or post-hoc control. Objective-specific policy prevents a point-only benchmark result from being presented as probabilistic customer-service performance.

No direct backtest placebo is available. A strong design would run the frozen production artifact in replay using only source objects received before each issue clock, then compare those outputs with immutable issued curves. It would include a latest-state-data placebo to measure how much apparent backtest skill comes from restated inputs. None of those measurements appears here.

The family protocol calls for blocked uncertainty for inferential claims, while this descriptive result reports no interval. Serial dependence remains unquantified. No unadjusted inferential claim or Holm-adjusted discovery is made.

## Limitations

The title promises a direct contrast that the outcome does not contain. The paper can explain governance mechanisms and report live performance, but it cannot quantify backtest-live divergence. A later study must freeze both arms before making that claim.

Pooled mean MAE can hide zone, horizon, and version heterogeneity. The served model can change over time. Persistence can also have different difficulty across regimes. Without per-cell paired differences, the headline may reflect differing availability of baseline rows.

Live lineage proves which model was selected, not necessarily that all source receipts, feature tensors, artifact hashes, and output curves are available for byte-for-byte replay. A complete operational explanation requires those identities. Backtests additionally face historical source-vintage limitations: a current database value may not equal what was knowable at the old issue time.

Operational scoring can also differ from a development evaluation in row eligibility. Missing forecasts, incomplete curves, late issues, and unavailable baselines may be rejected or represented differently. If a backtest silently evaluates only rows for which every feature is present, while production must issue or fall back on difficult days, the samples answer different questions. This paper's aggregate does not quantify that selection mechanism.

Model replacement adds another seam. A live record can contain successive champions chosen under evolving policies, whereas a backtest often evaluates one frozen family over the whole window. Pooling either side without version and objective identities can turn an operational history into a fictitious single model.

## Practical implication

Users should trust prospective scorecards more than unaudited retrospective claims when evaluating an operational forecast. A backtest remains useful for development, but it should be labeled with its information clock, artifact, selection process, and comparator. Live scores should bind to the actual served version.

The 0.7614007042253661 EUR/MWh improvement shows a positive pooled difference over persistence in MAE terms, but it is small enough that local zone and horizon evidence remains essential. It says nothing about probabilistic calibration or customer savings. The [Voltcast Research Content Plan](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/RESEARCH-CONTENT-PLAN.md) requires live losses and model-version distinctions to remain visible rather than being averaged into a marketing-only backtest number.

For reviewers, the minimum acceptable comparison is a common target panel with explicit missingness and separate labels for reconstructed, shadow, externally submitted, and customer-served curves. A single “accuracy” field without those identities is not enough to explain a backtest-live gap.

## Reproducibility

Verify the evidence and manifest identities and all four provenance hashes. Join accuracy to serving lineage on zone, delivery date, horizon bucket, and exact model version. Retain non-null served MAE and collect non-null baseline MAE. Calculate their pooled means and subtract served mean from baseline mean.

A direct extension should preregister a paired backtest-live panel, freeze one artifact and objective, replay point-in-time source vintages, bind every curve and source hash, report missingness, and block-resample complete delivery dates. Latest-state and wrong-version placebos should be reported separately. The public page follows the assigned Google Search Central [Article structured data](https://developers.google.com/search/docs/appearance/structured-data/article) source for machine-readable title, date, image, and evidence linkage; markup is not accuracy evidence.

## Disclosure

Analysis and drafting were model-assisted. The language model explained why backtests and live scores answer different questions and preserved the absence of a direct paired backtest result. It did not supply empirical values. Sources, code, assumptions, limitations, and hashes are disclosed. This working paper is not peer reviewed.

Volt has zero live traders and zero live capital. C0R is the only paper strategy. Production forecasting is a non-trading service. This accuracy analysis is not trading advice, financial advice, a customer-savings estimate, or authorization for model promotion.

## References

1. Voltcast. “Objective-specific seven-day forecast promotion.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/OBJECTIVE-PROMOTION-POLICY.md
2. Voltcast. “Volt 6.1 cross-zone price model card.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/V6.1-MODEL-CARD.md
3. Voltcast. “Voltcast Research Content Plan.” https://github.com/ossedk/voltcast/blob/main/docs/voltcast/RESEARCH-CONTENT-PLAN.md
4. Google Search Central. “Article structured data.” https://developers.google.com/search/docs/appearance/structured-data/article
5. ENTSO-E. “Single Day-ahead Coupling (SDAC).” https://www.entsoe.eu/network_codes/cacm/implementation/sdac/
