A backtest should ask what a system could have predicted at a historical moment using information available then and a model permitted to update then. Putting older dates in training and newer dates in testing does not fully answer that question. This article addresses enterprise load, demand and production forecasting, focusing on the time semantics of evaluation. All business dates, windows and numbers are constructed examples; no model was trained or experimentally tested.
Fix the prediction time, target range and update schedule
Suppose a factory publishes electricity forecasts every day at 08:00 for the next seven complete calendar days, beginning at 00:00 the following day, to support scheduling. The model retrains each Monday at 06:00 and enters service at 08:00, with no retraining during the day. Fix the business timezone and handle daylight-saving and other calendar differences according to site rules. Retain publication date, target date and model version.
A backtest that retrains daily evaluates a different service from weekly production retraining. Evaluating tomorrow alone also does not establish performance for seven-day scheduling. Specify forecast origins, horizons, publication and training frequencies, training cutoffs and the latest availability time of each feature.
Rolling-origin evaluation repeats forecasts at multiple historical origins. Forecasting: Principles and Practice describes constructing predictions from observations preceding the test observation and extending evaluation to multiple steps where needed.[1] This article additionally requires replaying data arrival and model updates, rather than automatically refitting at every test point.
Data have at least two clocks; labels also have a maturity time
Event time answers which day a production figure describes; availability time answers when the prediction service could actually read it. If September 8 production first arrives at 10:00 on September 9, an 08:00 forecast that day cannot use it even though its business date is in the past. Preserve revisions and the availability time of each version.
Feature construction should constrain both business time and availability, selecting an appropriate version available before the origin. Feast documents historical feature lookup and explicitly warns that constraining event time alone can still return later backfills or corrections.[3] A “point-in-time join” label therefore does not by itself establish faithful historical availability.
Prefer the time at which the service could actually read a value, rather than an upstream claim of publication. Lake ingestion, cleaning and feature computation can add delays. If historical snapshots cannot be recovered, disclose that limitation, conservatively simulate delay or exclude the field from strict backtesting. Today’s fully revised table must not masquerade as the historical serving view.
Label maturity determines when a training example becomes eligible. If daily electricity usage is reconciled two days later, it can subsequently be used for scoring but cannot become a known label at an earlier training time. A seven-day-total target becomes trainable only after the entire target interval ends and the agreed maturity conditions are met.
Derive splitting gaps from time semantics
scikit-learn’s TimeSeriesSplit supplies a gap that excludes samples between the end of training and testing. Its documentation also identifies equally spaced samples as a condition for folds covering comparable durations.[2] The unit is samples: the splitter does not automatically understand weeks, label delay or device groups.
First check whether each training label had matured by the simulated retraining cutoff, and whether its inputs were available at that example’s original forecast time. Regularly spaced data may allow conversion to a fixed gap. Missing observations, multiple devices or variable delays favor timestamp- and maturity-based masks. Setting gap to seven is not automatically correct for a seven-day target.
Manage a seven-day historical feature window separately from a seven-day future target window. Test examples using genuinely available observations from the training period can reflect legitimate serving behavior; shared history alone does not require removal. Exclude future observations, test targets and labels still unknown across the training boundary. If the objective is generalization to new devices, add device isolation; time splitting cannot establish it.
A larger gap reduces training data and may simulate a stale model that production would never use. Explain the information path it blocks. Randomly shuffled rows may serve as a diagnostic comparison but are not formal acceptance evidence for the future-forecasting task considered here.
Replay the serving schedule and retain the evidence
First freeze the evaluation origins, seven-day horizon and scoring rules. Cover normal production, shutdowns, holidays and seasonal changes, reserving a final period untouched by model selection. If history lacks a condition, report it as untested instead of manufacturing representativeness.
Second, at each simulated Monday 06:00 retraining time, select examples using historical visibility and maturity. Fit imputation, scaling, feature selection and the model only on those examples. scikit-learn recommends learning preprocessing from training data and using Pipeline to reduce misuse.[4] Pipeline does not repair an input table that already contains future information.
Third, reconstruct the feature snapshot at 08:00 each day and use the latest completed model allowed into service. If assumed training cannot finish by 08:00, retain the old model or record an unavailable-model state; do not give a backtest unlimited free computation. Save the model version, input snapshot, prediction for each target date and failure reasons.
Fourth, future exogenous variables such as temperature must use forecasts published at the time, not later observed weather. Calendars known in advance are usable; production plans require the version approved at that origin. If serving depends on another model predicting a feature, include that upstream prediction and its errors in evaluation.
Fifth, join actual outcomes once scoring conditions are satisfied. Backfilling labels for evaluation differs from backfilling features to generate historical predictions. Saved snapshots should reproduce the same inputs on rerun. Record seeds, versions and execution conditions for stochastic software, while recognizing that a fixed seed alone is not a complete reproducibility guarantee.
Constructed counterexamples: three plausible backtests overstate benefits
Consider the daily 08:00 seven-day forecast again. Incorrect approach A fills every historical feature from today’s final production table using business dates. This lets the September 9 forecast see data that did not arrive until 10:00. Correct it with availability-time filtering, retaining the missing state as it actually existed.
Incorrect approach B retrains at every daily test origin using all labels whose business dates precede that day. It ignores both weekly retraining and the two-day reconciliation delay. Correctly replayed, Wednesday uses Monday’s model, and Monday’s training excludes labels that were immature then. Differences cannot be attributed solely to algorithm quality.
Incorrect approach C predicts the first day, then uses its subsequently observed actual usage to predict the second, reporting all seven outputs as one seven-day forecast. This simulates receiving new information daily. If production publishes seven days at once, recursive forecasts must use prior predicted values. If it updates every day, store each update as a separate origin’s forecast.
These errors can make complex models appear better than simple baselines and can change candidate rankings. This article does not measure the size of that effect. Repair information boundaries before comparing models; a decline in one corrected backtest does not establish the same defect across all historical projects.
Report errors by horizon and business consequence
For each of the seven target days, report MAE, the mean absolute error for that horizon, in original units such as kWh. RMSE is the square root of mean squared error and is more sensitive to large errors. Include sample counts, date ranges and missing-data rules. Percentage errors can become unhelpful when shutdowns push actual values close to zero.
Include simple seasonal baselines such as the corresponding historical weekday, evaluated at identical origins with identical availability conditions. Training may use expanding history or a fixed rolling window: the former retains more data; the latter reduces the influence of older operating regimes. Choose window length and hyperparameters in development backtests, not by repeatedly consulting the final test period.
Daily seven-day forecasts predict the same target date several times. Keep both origin and horizon and compare by horizon first. Define weights beforehand if aggregating, so repeated targets do not obscure distant-horizon performance. Errors are correlated; treating every prediction as independent does not justify narrow confidence intervals.
If overprocurement and undersupply matter differently, business owners should define their costs and the corresponding loss. Lowest MAE does not automatically mean greatest business benefit. Separately report coverage for dates with missing forecasts, stale inputs or late publication, and include them in process evaluation under the agreed rules.
Deliver a backtesting protocol, not just a score
A minimum deliverable includes prediction and retraining schedules, feature-visibility rules, label-maturity rules, fold data manifests, model and preprocessing versions, per-origin predictions, baseline results and errors by operating condition. Agree acceptance criteria in advance, including tolerated regressions and fallback baselines for missing inputs. This article proposes no universal numerical threshold as an industry standard.
Before activation, run shadow forecasts against the real clock, recording inputs and outputs, then compare them with backtest assumptions. If data consistently arrive late, retraining misses its deadline or plan revisions cannot be traced, repair data and execution processes first. This connects enterprise data engineering and model evaluation into an explainable delivery process and creates a common basis for future model replacement.
References
- [1] Hyndman and Athanasopoulos: Forecasting: Principles and Practice, 3rd ed., Time series cross-validation
- [2] scikit-learn: TimeSeriesSplit (stable documentation, accessed 2026-10-03)
- [3] Feast: Point-in-time joins (accessed 2026-10-03)
- [4] scikit-learn: Common pitfalls and recommended practices (accessed 2026-10-03)
