---
id: VOLT-HOME-WP-037
title: "How much battery value is lost when forecasts replace perfect foresight?"
slug: how-much-battery-value-is-lost-when-forecasts-replace-perfect-foresight
description: "A serving-lineage MAE sensitivity reaches a 100% haircut relative to captured spread, establishing a bound rather than a measured dispatch loss."
published: 2026-08-30
cluster: "Home batteries"
status: measured
evidence_url: /research-data/home-papers/how-much-battery-value-is-lost-when-forecasts-replace-perfect-foresight.json
figure_url: /research-media/home-papers/how-much-battery-value-is-lost-when-forecasts-replace-perfect-foresight.webp
figure_alt: "Chart for How much battery value is lost when forecasts replace perfect foresight?: forecast-error haircut relative to captured spread, shown as spread, MAE."
source_ids:
  - energy-informatics-storage
  - acer-retail-2025
  - iea-electricity-2026
  - dynamic-tariff-viability
  - volt-research-content
peer_reviewed: false
---
## Abstract

The frozen evidence reports a forecast-error haircut relative to captured spread of 100%, based on 17,750 observations. The regenerated JSON reports no interval for this estimand. The evidence interpretation is essential: this ratio is a sensitivity bound using serving-lineage mean absolute error, not a measured battery dispatch loss. A 100% point estimate means the error magnitude used in the comparison is as large as the captured spread on the metric’s scale; it does not mean every forecast-based schedule loses all BESS-v2 value.

This paper uses the result to explain why perfect foresight must be separated from deployable scheduling. BESS-v2 selects intervals from the realized day-ahead curve and is therefore a ceiling-like benchmark. Forecast replacement can change both which intervals are chosen and the spread ultimately settled. Measuring actual loss requires a frozen issue-time forecast, a dispatch policy, identical device constraints, and realized settlement. The evidence provides an error-to-spread sensitivity ratio, not that end-to-end counterfactual. Wholesale-only accounting and normalized per-MW conventions further limit translation to household economics.

## Plain-language answer

The reported sensitivity reaches 100%. In plain terms, the serving-lineage forecast error used by the analysis is as large as the indexed spread used for comparison. That is a warning that a perfect-hindsight battery value can be entirely consumed on a simple error-magnitude basis. It is not proof that actual forecast dispatch captures zero.

Forecast error does not translate mechanically into dispatch loss. An error can occur in an interval the battery never uses, shift both charge and discharge prices together, or leave their rank unchanged. Conversely, a smaller error can swap the cheapest and most expensive choices and cause a large loss. To know the real haircut, the schedule must be selected from the forecast and settled against realized prices. The current result is therefore a conservative sensitivity bound. It says perfect-foresight values should not be presented as attainable without a forecast-based replay.

## Research question

The causal counterfactual of interest is the difference between two schedules on the same delivery curve: one selected with realized prices and one selected with a forecast frozen before the decision. Both should obey identical duration, power, efficiency, cycle, state-of-charge, degradation, and tariff rules. Their settled value difference, divided by the perfect-foresight value under a declared convention, would measure realized forecast replacement loss.

The paper-level metric instead compares serving-lineage MAE with captured spread. MAE is a point-forecast accuracy statistic. It summarizes absolute price error but does not preserve interval ranks or dispatch choices. The research question is thus only partially answered. The ratio identifies whether forecast error is economically material relative to the spread, while the actual schedule haircut remains unmeasured in this evidence object.

## Data and provenance

The analysis is tied to SELECT-only snapshot SHA-256 `7e97489fc8528c8cc8c38830e05b48d949ce1f67b98575dff26f5d7c321e4c67`, with a read-only transaction and 180-second timeout. Registered source contracts are `day_ahead_prices`, `bess_index_daily`, `bess_forecast_daily`, `generation_mix`, and `capture_stats`. Daily prices span 2021-01-01 to 2026-08-29, detailed intervals span 2025-10-01 to 2026-08-29, and the long-history boundary spans 2015-01-01 to 2026-08-29. Publication cut-off is 2026-08-30T00:00:00Z.

Unlike the other battery papers in this block, this metric has sample size 17,750. The unit is percent. Analysis code SHA-256 is `57c57de79cdab2b5b6d6c54c485cb5162598c5ba0b0bfe995da40d75e6c52ba9`; protocol, `adb36bf6b447af9f96339249b8becaefc20422499cca1977242866347a97bd4b`; registry, `7bcb91d7476d0a69fe9fa75a5c7782f8117e0153f82f9112b7e1d307d3943717`; source registry, `07949550ac443ff673fda5c0209b99f137544f3ffecf6775f109bb9d09663bd6`. Those hashes bind the sensitivity to the serving lineage and corpus available before the publication cut-off.

## Method

BESS-v2 is normalized per unit of power and uses perfect foresight of realized day-ahead prices. That perfect-information convention identifies an ex post opportunity ceiling under the index rules. It does not claim a live controller had those prices when choosing a schedule. The forecast side contributes serving-lineage MAE, preserving the provenance of forecasts actually designated by the system rather than selecting a favorable model after observing outcomes.

The primary ratio expresses forecast error relative to captured spread and caps or reports the haircut at 100% in the frozen output. This is a sensitivity comparison, not a schedule simulation. Wholesale-energy accounting excludes retail taxes and network charges. No unadjusted significance claim is made; within-family Holm control applies to inferential tests. No interval is reported, and the sensitivity cannot be converted into measured dispatch loss.

## Results

The primary value is 100%. The point estimate says the forecast-error magnitude is equal to the captured spread on the chosen scale. The figure publishes spread = 63.11932374144673 and MAE = 87.14137855774648 on their displayed scales. Their comparison motivates the capped sensitivity, but neither figure value is a measured dispatch loss. The evidence interpretation calls the ratio a sensitivity bound, and no uncertainty interval is reported.

This finding materially limits claims based on perfect foresight. If model error is of the same order as the spread, a forecast schedule may frequently mis-rank valuable intervals or capture less spread. But the exact loss could be smaller or larger for a given schedule because MAE ignores rank and covariance. No settled forecast-dispatch value, capture percentage, negative-cycle count, or zone-level result is published here. The answer is therefore not “exactly all value is lost”; it is “the published sensitivity is large enough to consume all captured spread, so direct replay is necessary.”

## Robustness and placebo checks

The strongest follow-up is an end-to-end matched replay. Freeze each price forecast at its issue time, optimize the battery without realized prices, settle the selected intervals against the realized curve, and compare with the same asset under perfect foresight. Include persistence and simple rank heuristics as controls. Report capture ratio, negative-settlement frequency, and sensitivity to degradation and efficiency.

Placebo tests should shift forecast issues to the wrong delivery day, randomize interval ranks, or use an unconditional daily profile. A serving model should beat those controls on an economic objective before a “safer” or “higher value” claim. Results should also be stratified by spread width: high-MAE days may still be usable when spreads are extreme, while small-spread days can be fragile. None of these dispatch checks is in the current JSON, and no uncertainty range is reported for the sensitivity ratio.

## Limitations

MAE is not an economic loss function. It weights all intervals equally, even though a battery uses only selected charge and discharge periods. It does not preserve signs, ranks, joint errors, or temporal dependence. Dividing MAE by captured spread can be unstable when spreads are small and may mix zones or durations. The paper-level evidence does not describe those distributional details.

Perfect foresight is unattainable at the scheduling clock. BESS-v2 is normalized and wholesale-only, excluding retail taxes, network charges, degradation, auxiliary use, and contract-specific import-export terms. The 100% ratio is not observed household cash and does not establish that every schedule loses its value. The evidence also does not publish a forecast policy, issue-horizon breakdown, or direct perfect-versus-forecast pair. These limits make the result a strong warning and a weak point estimate of actual dispatch loss.

The ratio should also be separated by decision horizon. A forecast issued close to delivery may rank intervals better than a multi-day forecast, while a day-ahead auction decision has a fixed information cut-off. Pooling horizons can obscure that distinction. Any replay must pin the exact issue used for each target interval and reject forecasts published after the declared clock. It should never backfill a later, more accurate model into an earlier schedule.

Economic scoring should preserve the paired nature of charge and discharge. Subtracting average error from average spread ignores whether errors on the two selected legs cancel or compound. A paired replay can report selected-versus-oracle interval overlap, rank regret, captured spread, and settled value while keeping the normalized asset constant. Those outputs would explain why the 100% sensitivity does or does not translate into an equally large schedule haircut. They are not available in the frozen evidence and remain unresolved.

## Practical implication

Whenever a BESS index uses realized prices, label it as a perfect-foresight ceiling and publish a forecast-based capture measure separately. Do not discount the ceiling with a generic MAE percentage and call the remainder attainable. Instead, select the schedule using the exact forecast available at the decision clock, then settle it on actual prices. Preserve model lineage so a later winner cannot replace the model that was actually available.

For a controller, focus on rank and spread uncertainty around candidate charge and discharge intervals. A policy can require conservative separation between forecast quantiles, decline cycles when rank confidence is low, and maintain reserve. The 100% sensitivity indicates that forecast quality is not a minor correction to the indexed value. It is a first-order constraint that should be visible in every practical comparison.

## Reproducibility

The public evidence is `/research-data/home-papers/how-much-battery-value-is-lost-when-forecasts-replace-perfect-foresight.json`. Verify status `measured`, sample size 17,750, value 100%, absence of an interval, figure values, and the sensitivity-bound interpretation. The evidence SHA-256 is `a04e69bcb1fac3b9f779f37b153f999751a3e7ffeba6c5b92a60611648be09fc`; figure SHA-256 is `0032b55cb6d655e0dde3ddee017837b001f0ac76861c28dad89ff0b50548dba8`.

Reproduce with the exact snapshot, analysis code, protocol, registry, windows, cut-off, and serving lineage. Keep MAE and captured spread on the registered scales. Label the output as a sensitivity ratio. A direct dispatch-loss extension requires frozen forecasts, schedule code, asset assumptions, and realized settlement rows. Licensing is governed by `/legal/data-licensing` and [Voltcast data licensing and redistribution](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/LICENSING.md).

## Disclosure

Analysis and drafting were model-assisted. The result is deliberately described as a sensitivity bound, following the evidence interpretation. This working paper is not peer reviewed and is not trading advice, financial advice, an investment recommendation, or operational authorization. Voltcast has no live traders or live capital. No capital or battery action is implied by a forecast-error ratio.

## References

- `energy-informatics-storage` — Energy Informatics. [Risk and reward: evaluating household energy storage for optimizing demand-side flexibility under dynamic tariffs](https://doi.org/10.1186/s42162-025-00602-9). Kind: peer-reviewed.
- `acer-retail-2025` — ACER and CEER. [Rewarding flexibility: How retail contract choice can help unlock consumer flexibility](https://www.ceer.eu/wp-content/uploads/2025/11/ACER-CEER-2025-Retail-monitoring.pdf). Kind: official.
- `iea-electricity-2026` — International Energy Agency. [Electricity 2026](https://www.iea.org/reports/electricity-2026). Kind: official.
- `dynamic-tariff-viability` — Advances in Applied Energy. [Assessing the conditions for economic viability of dynamic electricity retail tariffs for households](https://doi.org/10.1016/j.adapen.2024.100174). Kind: peer-reviewed.
- `volt-research-content` — Voltcast. [Voltcast Research Content Plan](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/RESEARCH-CONTENT-PLAN.md). Kind: canonical.
