Why a missing ENTSO-E A03 point is not missing data
Why a missing ENTSO-E A03 point is not missing data. Interval counts alone confuse a valid local-clock transition with missing delivery energy.
Why a missing ENTSO-E A03 point is not missing data
Abstract
Market-data quality checks often treat an unexpected interval count as proof that a price point is missing. That rule fails when local delivery days legitimately change length. This paper evaluates naïve missing-data flags against a clock-safe completeness test based on duration and contiguous UTC timestamps, not document code alone. Among 948 naïve flags, 16 are recorded as valid daylight-saving-time-length cases and 932 remain in the figure's other flag category. The primary artifact reports 1.6877637130801688% of naïve flags explained by the DST-valid cases. It reports no interval for this estimand.
The finding is not that missing data never occurs, nor that an ENTSO-E-labelled record can be accepted without validation. It is that row counts and a document-code expectation are insufficient evidence. Physical interval continuity, duration, zone clock, and revision provenance must be checked together. The detailed analysis covers ten representative European bidding zones from 2025-10-01. It is not a household bill study. Analysis and drafting were model-assisted; the paper is not peer reviewed and is not trading advice.
Plain-language answer
A price file can look one point short while still covering the entire local delivery day correctly. On a daylight-saving transition, the local clock can skip or repeat part of the ordinary label sequence. If a validator expects every date to have the ordinary number of labels, it can call a valid day incomplete. That false alarm may then trigger a dangerous repair: filling an interval that never existed, shifting all later prices, or rejecting a usable schedule.
The evidence found 16 such valid cases among 948 naïve flags. Most naïve flags were therefore not explained by this specific condition, so the correct lesson is not “ignore missing-point alerts.” The lesson is “verify them with the timeline.” A missing label, a missing physical interval, an upstream omission, and a legitimate short or long local day are different states.
For household automation, that distinction matters before optimization starts. An EV charger or heat pump needs a sequence of physical delivery intervals. If software invents a point to make an array look normal, the resulting schedule may associate a command with the wrong price or time. If software rejects every nonstandard day, it may fall back unnecessarily. A sound completeness test uses explicit interval boundaries and timezone identity, then reports the reason for any unresolved gap.
Research question
The research question is how often a naïve missing-data flag in the registered interval sample is explained by a valid daylight-saving-time-length day. The estimand is the percentage of naïve flags attributable to that clock condition, using duration and contiguous UTC timestamps as the completeness basis.
This paper does not test all causes of missing records. It does not evaluate the reliability of ENTSO-E as a whole, compare vendors, or infer the quality of a household integration. It tests one precise failure mode in validation logic: confusing a non-ordinary local-day shape with absent delivery energy.
Data and provenance
The evidence lists daily prices from 2021-01-01 through 2026-08-29, detailed intervals from 2025-10-01 through 2026-08-29, and long history from 2015-01-01 through 2026-08-29. The limitation states that detailed interval analysis uses ten representative European bidding zones from 2025-10-01. The primary sample comprises 948 naïve flags rather than all price intervals or all delivery days.
The source contract includes zones, day_ahead_prices, grid_revisions, auction_publications, and ingestion_runs. Zone records anchor local-time interpretation. Price rows provide interval boundaries and values. Revision and publication records preserve lineage, while ingestion records distinguish data-state observations from market facts.
The publication cutoff is 2026-08-30T00:00:00Z. The query ran in a read-only transaction with a 180-second statement timeout. The frozen snapshot SHA-256 is 7e97489fc8528c8cc8c38830e05b48d949ce1f67b98575dff26f5d7c321e4c67. Analysis code is identified by 57c57de79cdab2b5b6d6c54c485cb5162598c5ba0b0bfe995da40d75e6c52ba9, the protocol by adb36bf6b447af9f96339249b8becaefc20422499cca1977242866347a97bd4b, the paper registry by 7bcb91d7476d0a69fe9fa75a5c7782f8117e0153f82f9112b7e1d307d3943717, and the source registry by 07949550ac443ff673fda5c0209b99f137544f3ffecf6775f109bb9d09663bd6. The evidence manifest records evidence SHA-256 b046405855fb4b7ea7b2482a8c3048fff2d81245f785b85ffadfefb32834f291 and figure SHA-256 c2183bd168c9c19f4da7f8a7c0d0e57f9527f27d046002fa80abc88f910b8a97.
Those identities permit provenance checking. The public artifact exposes aggregates and methodology rather than the underlying licensed payload. Reproduction must follow the referenced licensing posture and should not infer redistribution permission from publication of these counts.
Method
The method is part of the registered paired interval reconstruction, clock-safe counterfactual, and day-block bootstrap family. It starts with a naïve detector that flags a day when its point pattern does not match the ordinary expectation. The clock-safe adjudication then examines duration and contiguous UTC timestamps. The evidence explicitly assumes that completeness is assessed by those physical properties, not by document code alone.
This ordering prevents a circular result. The initial flag represents what a simplistic consumer might report. The adjudication asks whether the allegedly missing point corresponds to an actual gap in physical delivery coverage. When the timestamps remain contiguous and the total duration matches a legitimate local-day transition, the case is classified as DST-valid rather than repaired as missing.
Curve type A03 adds a separate reason that a missing encoded position need not mean missing delivery data. In the post-seam ENTSO-E A44 representation, an unchanged price may be omitted because the preceding value continues until the next change. A parser must expand that stepwise representation onto the complete native MTU grid before it tests interval coverage. It must then evaluate unambiguous UTC continuity and the bidding zone's legitimate 92-, 96-, or 100-quarter-hour local-day shape; blindly inserting a null or treating every absent encoded point as a gap is incorrect.
The numerator is the count of naïve flags explained by valid DST-length days; the denominator is the full flagged sample. The secondary aggregate records the numerator as 16 zone-days. The primary sample size is 948 flags, and the reported percentage is 1.6877637130801688.
The regenerated artifact stores bootstrap_95_interval as null and labels its interval method not reported for this estimand. Within-family Holm control applies to inferential claims, but this result is explicitly descriptive and makes no unadjusted significance claim.
Results
Of 948 naïve missing-data flags, 16 are classified as DST-valid zone-days by the duration-and-contiguity test. The primary result is 1.6877637130801688% of naïve flags explained by DST-length days. The figure labels DST-valid and other flag carry exact values 16 and 932. This directly supports the limited interpretation in the evidence: interval counts alone can confuse a valid local-clock transition with missing delivery energy.
The result also puts a boundary on the claim. DST-valid cases are a minority of the flagged sample. The remaining flags cannot be declared real gaps from this paper, because the evidence does not publish their adjudicated causes. They remain “other flag” in the figure rather than being relabelled as upstream failures.
No confidence interval is reported in the frozen JSON. The descriptive percentage must not be paired with the stale 1.0 endpoints from an earlier artifact, nor with an interval reverse-engineered from the counts.
Robustness and placebo checks
The strongest check is the use of contiguous UTC timestamps and weighted duration against the naïve interval-count flag. UTC or another unambiguous instant distinguishes physical continuity from repeated or skipped local labels. Duration checks whether the covered energy-delivery span is coherent. Together they form a clock placebo: if a flag disappears only because the local day has a legitimate transition shape, it should not be treated as absent energy.
The method does not accept every nonstandard day automatically. Zone identity is required to interpret its local clock, and physical timestamps must remain contiguous. A malformed sequence cannot pass merely by carrying a familiar document code. This guards against the opposite error of using DST as a blanket excuse for corruption.
The family protocol permits blocked uncertainty procedures where estimator-valid, but this estimand's current evidence reports no interval. Within-family Holm control prevents descriptive outputs across the family from being promoted through selective unadjusted tests. No causal claim is made.
The null interval is preserved rather than filled from an obsolete serialization. Any future uncertainty estimate should be issued through a newly governed evidence artifact, not an undocumented edit to this paper.
Limitations
Detailed intervals come from ten representative European bidding zones beginning on 2025-10-01. That sample does not establish the prevalence of false flags across all ENTSO-E areas, years, parsers, or document types. Clock rules and source formats can differ.
The paper tests one explanation for naïve flags. It does not publish a taxonomy of the remaining flagged cases. No conclusion is made about whether those observations are missing, revised, duplicated, delayed, or filtered for another reason.
The evidence does not observe household commands, state of charge, indoor temperature, comfort, or bills. This is not a household bill study and cannot quantify consumer loss. It identifies a data-integrity condition upstream of those outcomes.
Finally, no uncertainty interval is available for the percentage point estimate. The primary count and percentage remain reportable as frozen descriptive aggregates, while inferential use is unsupported.
Practical implication
A production validator should distinguish at least four questions: Has an A03 step curve been expanded at its declared native MTU? Is the local label pattern ordinary? Are physical timestamps contiguous? Does total interval duration cover the legitimate delivery day for the zone? Only the expanded physical timeline can establish delivery completeness. A document code or raw encoded-point count can be useful routing metadata, but it should not be the final verdict.
When a check fails, automation should retain the original data and attach a reasoned status. A DST-valid day can proceed through a timezone-aware scheduler. A real physical gap should trigger a bounded fallback or halt. An ambiguous day should remain unresolved rather than being “fixed” by duplicating, interpolating, or shifting prices without authority.
For home systems, the controller should consume explicit start and end instants and conserve requested energy across those durations. Tests should include legitimate short and long local days, repeated labels with different offsets, and truly missing physical intervals. Success means the controller treats those cases differently.
Reproducibility
The public evidence entry is /research-data/home-papers/why-a-missing-entso-e-a03-point-is-not-missing-data.json. It contains the frozen assumptions, windows, source tables, result, affected count, limitation, disclosure, and provenance identities. The figure URL and exact alternative text are included in frontmatter.
A reproduction should verify all hashes before querying. It should preserve the cutoff, read-only transaction, and detailed sample definition. First run the naïve point-pattern check. Then, without altering records, reconstruct each flagged zone-day from unambiguous interval starts and ends, verify contiguity, calculate weighted duration, and apply the zone's local-clock interpretation. Aggregate DST-valid adjudications over the original flagged denominator.
Reviewers should compare the primary aggregate, the exact figure vector [16, 932], and the null interval to the public JSON. If implementation results differ, investigate timezone data, interval boundary conventions, eligibility filters, and artifact serialization before changing any claim.
Disclosure
Analysis and drafting were model-assisted; sources, code, assumptions, and evidence hashes are disclosed. This working paper is not peer reviewed, not a household bill study, and not trading advice. Volt has zero live traders and zero live trading capital; C0R is the only paper strategy. Production weather forecasting is a non-trading service. The data-quality result authorizes no order, capital, or live-trading action.
References
- Spam Policies for Google Web Search — Scaled content abuse — Google Search Central.
- Single Day-ahead Coupling (SDAC) — ENTSO-E.
- EU electricity trading in the day-ahead markets becomes more dynamic — European Commission.
- ACER Decision 13-2024 on SDAC Products — Agency for the Cooperation of Energy Regulators.
- Voltcast Architecture — Voltcast.