---
id: "VOLT-HOME-WP-097"
title: "How stale can a forecast become before automation should reject it?"
slug: "how-stale-can-a-forecast-become-before-automation-should-reject-it"
description: "Why freshness must be tied to issue and target windows; the measured receipt-age statistic does not define a universal rejection threshold."
published: "2026-08-30"
cluster: "Reliable automation and machine-readable energy evidence"
status: "measured"
evidence: "/research-data/home-papers/how-stale-can-a-forecast-become-before-automation-should-reject-it.json"
figure: "/research-media/home-papers/how-stale-can-a-forecast-become-before-automation-should-reject-it.webp"
figure_alt: "Chart for How stale can a forecast become before automation should reject it?: median auction-receipt age from prior UTC midnight, shown as P10, Median, P90."
source_ids:
  - "google-article"
  - "google-structured-data"
  - "google-scaled-content"
  - "entsoe-sdac"
  - "volt-architecture"
peer_reviewed: false
---

## Abstract

Freshness is not the age of an HTTP cache entry in isolation. A controller must relate issue time, publication detection, target delivery interval, source revision, and the moment a household action becomes irreversible. The frozen evidence measures a receipt-clock statistic: across 2,141 production-origin auction publication receipts, the median detection time was 11.217222222222222 hours after the previous UTC midnight. The regenerated artifact reports no uncertainty interval for the median.

This result does not measure forecast degradation with age and does not establish a universal rejection threshold. Indeed, “11.217 hours old” would be a misleading description: the statistic is measured from a calendar anchor to auction detection, not from forecast issuance to use. The defensible integration rule is semantic freshness. Validate that the payload was issued or detected under the expected publication cycle, covers the exact future target window, has not been superseded, and leaves enough time for a safe action. Reject cached data that cannot prove those properties, even if recently fetched. Production receipts ground the timing distribution. Suggested frozen-clock, wrong-target, and superseded-payload scenarios are deterministic fault injections for future adapter tests, not repository tests claimed as measured.

## Plain-language answer

There is no evidence-grounded answer such as “reject every forecast after a fixed number of hours.” This study did not compare forecast error at different payload ages. It measured when auction publication receipts occurred relative to the previous UTC midnight. The median across 2,141 receipts was about 11.217 hours after that anchor.

A payload fetched seconds ago can still be stale if it describes yesterday, came from an older forecast issue, or was superseded. A payload stored for several hours can remain valid if it is the latest authorized issue and still covers the intended future interval. The controller should therefore validate the issue identity, target window, source revision, and decision deadline rather than rely on cache age alone.

If any required clock is missing or inconsistent, reject the new automated plan and use a declared appliance-safe fallback. Do not silently convert “unknown freshness” into “fresh.”

## Research question

The research question is how stale a price forecast can become before home automation should reject it. To answer that directly, one would need forecast issues, use times, target horizons, lineage, and accuracy or decision-regret outcomes across increasing ages.

The frozen primary metric is related but not equivalent. It asks when complete auction publication receipts were detected relative to the previous UTC midnight for each delivery date. That clock can characterize publication timing and help define an expected issue cycle. It does not observe forecast age or accuracy decay.

The paper therefore separates two questions. The measured question is the receipt-age distribution under a fixed calendar anchor. The engineering question is which invariants establish that a payload remains eligible for a particular action. The acceptance criterion is fail-closed: a controller uses a payload only if it can prove the correct issue, target, revision, and remaining decision window. It never infers a universal validity duration from the median receipt clock.

## Data and provenance

The public evidence identifies `VOLT-HOME-WP-097`, schema `volt-home-paper-evidence-v1`, status `measured`, and publication cutoff 2026-08-30T00:00:00Z. The primary sample contains 2,141 auction publication receipts. The extraction selected zone, delivery date, period count, detection timestamp, notification timestamp, and notification delay. The analysis derived freshness from the detection timestamp and the prior UTC midnight associated with the delivery date.

The source snapshot came from a SELECT-only production transaction with a 180-second statement timeout. The snapshot SHA-256 is `7e97489fc8528c8cc8c38830e05b48d949ce1f67b98575dff26f5d7c321e4c67`. Analysis-code, protocol, paper-registry, and source-registry hashes are respectively `57c57de79cdab2b5b6d6c54c485cb5162598c5ba0b0bfe995da40d75e6c52ba9`, `adb36bf6b447af9f96339249b8becaefc20422499cca1977242866347a97bd4b`, `7bcb91d7476d0a69fe9fa75a5c7782f8117e0153f82f9112b7e1d307d3943717`, and `07949550ac443ff673fda5c0209b99f137544f3ffecf6775f109bb9d09663bd6`. The evidence-manifest hashes are `31170d7f10df869afb4b7d2066a1d8efc57f3c4b37a2639ed86d3896ea06d6ab` for the JSON and `cea5e15e0e17820f1dd1bd5cfe506f1fdb8b9b6330d53ae05242998319fa5e33` for the WebP figure.

These are production-origin publication receipts. They are not forecast issue rows, accuracy outcomes, customer cache logs, or device decisions. The family source contract also lists grid revisions, ingestion runs, public API schemas, and repository integration fixtures, but no fixture-run result is reported here. Wrong-target and stale-cache examples below are proposed deterministic checks, not observed incidents.

## Method

For each auction publication receipt, the analysis took the delivery date, subtracted one day, constructed midnight in UTC, and calculated the elapsed hours from that anchor to `detected_at`. Negative values, if any, were floored at zero. The primary outcome is the median of the resulting 2,141 ages.

The value is 11.217222222222222 hours. The regenerated evidence records `bootstrap_95_interval: null` and `interval_method: "not reported for this estimand"`. This paper therefore reports no uncertainty interval for the median.

An operational freshness validator should use a richer tuple: product and zone, issue or publication identity, issued or detected time, target start and end, source revision, received time, validation time, and the action deadline. Policies can then express maximum transport delay, latest allowed issue, and minimum lead time separately. A single “age in hours” field cannot capture those dimensions.

Three design patterns have different failure modes. Fetch-time time-to-live is simple but can bless old content retrieved recently. Issue-time time-to-live is better but ignores target alignment. Target-aware eligibility checks both the latest authorized issue and the specific delivery window; it is more complex but directly answers whether the data can support the action. The evidence favors using receipt timing as one check within the target-aware pattern, not as the sole threshold.

## Results

The primary result is a median auction-receipt age of 11.217222222222222 hours from the prior UTC midnight, based on 2,141 receipts. The interpretation in the evidence states that automation should validate issued-at, target window, and freshness rather than trust a cached payload indefinitely.

![Chart for How stale can a forecast become before automation should reject it?: median auction-receipt age from prior UTC midnight, shown as P10, Median, P90.](/research-media/home-papers/how-stale-can-a-forecast-become-before-automation-should-reject-it.webp)

The figure values are exactly 10.350555555555555, 11.217222222222222, and 13.350555555555555 hours for P10, median, and P90. The evidence reports no interval around the median.

The result characterizes the receipt clock under the registered anchor. It does not say that a forecast is acceptable for 11.217 hours, should be rejected after 11.217 hours, or has a particular error at that age. No forecast-accuracy table enters the primary calculation. The practical finding is the need for clock-aware validation, not a universal numerical timeout.

## Robustness and placebo checks

A reproduction should verify the UTC anchor, the delivery-date subtraction, timezone-aware parsing of `detected_at`, the zero floor, and receipt uniqueness. Replacing UTC midnight with local midnight would answer a different question and could change the distribution around daylight-saving transitions. It must not be called the same metric.

Future adapter fault injections should include a fresh fetch carrying yesterday’s target, the correct target from a superseded issue, a payload with no issue timestamp, a source revision that regresses, a valid issue received after the action deadline, and a host clock shifted forward or backward. These are recommended deterministic fixtures, not measured repository tests. Every ambiguous case should be rejected or quarantined.

Placebos are equally important. A payload can be several hours past issue yet still be the latest valid issue for a future target; the validator should accept it when all declared policy conditions pass. Conversely, a payload retrieved moments ago but targeting an elapsed interval should fail. Those paired cases demonstrate semantic freshness rather than a crude cache timer.

If multiple candidate freshness thresholds are compared against forecast error or regret in a future study, they should be preregistered and corrected for multiplicity. The present result is descriptive and makes no inferential claim.

## Limitations

The primary metric concerns auction publication receipts, while the title asks about forecast staleness. Auction detection time is not forecast issue age. No forecast performance, decision regret, or accuracy degradation is measured as a function of age.

No median interval is available in the frozen JSON, and this paper does not invent one.

Coverage can vary by zone and by the launch date of the receipt rail. The aggregate does not publish zone stratification or missing-event denominators. Receipt timing also stops at server detection and does not measure customer cache or device clocks.

The family limitation applies: infrastructure receipts do not measure a customer’s network or firmware. No repository integration-test pass is claimed. A freshness policy must also account for appliance constraints and cannot be certified by this timing statistic. The paper makes no savings or trading claim.

## Practical implication

Represent freshness as a validation result with reasons, not a lone time-to-live. Require the expected product, zone, target start and end, current issue or publication identity, source revision, complete interval coverage, and trusted timestamps. Compare those fields with a server-side or otherwise trusted clock and the appliance’s latest safe decision time.

Define separate limits for transport freshness and decision eligibility. A cached response can be reused only while its issue remains authorized and its target intervals remain future and complete. A later revision should supersede the earlier observation, while an out-of-order revision should be rejected. Store the validated payload hash with the generated schedule.

If the validator cannot prove freshness, do not issue a new price-driven command. Mark the schedule stale, retain it only for audit, and activate the declared local fallback. Do not reset the freshness timer merely because the same old payload was fetched again. Recovery requires a valid current identity, not just network success.

## Reproducibility

Verify the frontmatter against the registry, all five provenance hashes, and both evidence-manifest artifact hashes. For each of the 2,141 frozen auction receipts, construct the previous UTC midnight from the delivery date, calculate non-negative elapsed hours to detection, and take the median. Confirm 11.217222222222222 hours and that the evidence reports no interval. A new analysis that estimates median uncertainty would be a methodological extension and should receive a new evidence version.

A direct staleness study should preregister forecast issue lineage, target horizons, use times, accuracy or decision metrics, threshold candidates, appliance deadlines, and fixtures. It should report production receipts separately from deterministic clock-fault scenarios.

## Disclosure

Analysis and drafting were model-assisted. This is a public working paper. It is not peer reviewed. Evidence hashes, clocks, assumptions, and limitations are disclosed. Voltcast publishes the cited internal architecture and produced the frozen aggregate.

No forecast-age performance result, customer cache, device, or repository-test run was measured. Volt has no live traders or live capital; C0R is the only paper strategy, and production weather forecasting is a non-trading service. This paper is not trading advice, an API freshness guarantee, or a universal device-control threshold.

## References

1. Google Search Central, [Article structured data](https://developers.google.com/search/docs/appearance/structured-data/article).
2. Google Search Central, [Introduction to structured data markup in Google Search](https://developers.google.com/search/docs/appearance/structured-data/intro-structured-data).
3. Google Search Central, [Spam Policies for Google Web Search — Scaled content abuse](https://developers.google.com/search/docs/essentials/spam-policies).
4. ENTSO-E, [Single Day-ahead Coupling (SDAC)](https://www.entsoe.eu/network_codes/cacm/implementation/sdac/).
5. Voltcast, [Voltcast Architecture](https://github.com/ossedk/voltcast/blob/main/docs/voltcast/ARCHITECTURE.md).
