Define what is forecast and when it is needed
Specify the target, unit, time interval and forecast horizon. Predicting tomorrow’s unit sales for one store differs from predicting total demand over a supplier’s replenishment lead time. Record the forecast issue time and the operational decision it supports. A model that predicts the next hour is not validated for a four-week purchasing decision.
Separate forecast generation from anomaly detection and operational optimization. They may share data, but their outputs and evaluation criteria differ. Agree who reviews the forecast and what happens when it is unavailable before choosing an algorithm.
Build a point-in-time data contract
Record event time and the time each input became available. Lagged values and rolling statistics must use information known when the historical forecast would have been issued. Future realized weather cannot stand in for the weather forecast available at that time. Preserve planned promotions as they were known, rather than substituting a later revised plan.
Distinguish missing observations from actual zero activity. Investigate stockouts, store closures, returns and changes in product identifiers. Observed sales during a stockout may not represent unconstrained demand. Document how these cases are handled and whether the evaluation target is sales or estimated demand.
Compare a seasonal baseline with candidate models
Start with a reproducible baseline, such as the last observed value or the value from a comparable prior season. Use the same information and forecast horizon for each candidate. Add model complexity when the evaluation demonstrates a useful improvement relative to the additional operating cost.
The official scikit-learn lagged-feature example compares shuffled evaluation with time-based splits and finds the shuffled result overly optimistic for its forecasting task. Its bike-demand example illustrates evaluation design, not an accuracy promise for a retailer or energy system.
Backtest the actual forecast workflow
Evaluate at successive historical issue times, using only the data available at each cutoff. Reproduce the intended retraining and forecast-generation schedule. Score the full horizon needed by the business, including later steps if a model feeds its predictions back into a recursive forecast.
Keep the final evaluation period separate from model and feature selection. Review results across seasons, product groups, sparse series and unusual events. Record which series were excluded and why. Use any gap between training and evaluation according to feature availability and the task, rather than copying a tutorial setting without justification.
Report error in understandable units
Mean absolute error averages the absolute difference between forecast and observed values. In a hypothetical three-period example, observations of 10, 20 and 30 units with forecasts of 12, 18 and 27 produce absolute errors of 2, 2 and 3. MAE is 7 divided by 3, approximately 2.33 units per period. This arithmetic does not establish a percentage accuracy or a measured TensorBlue result.
Report systematic over- or underprediction and error by horizon as well as the aggregate. State how metrics combine series of different scales. Percentage errors need an explicit policy for zero or near-zero targets. If forecasts include intervals, evaluate their observed coverage and width on held-out periods; a wide interval can cover observations while being less useful for a decision.
Measure the decision and operate the release
A lower forecast error does not automatically establish fewer stockouts, lower energy costs or a better margin. Evaluate the downstream ordering or allocation policy with its constraints and cost assumptions, then monitor actual outcomes over an agreed pilot. Keep simulated results separate from observed business results.
Version the model, input snapshot, preprocessing, forecast issue time and horizon. Monitor data freshness, missing series, failed jobs and forecasts delivered too late to use. Provide an agreed baseline fallback and assign ownership for overrides, retraining and rollback. Budget data preparation, experiments, integration and recurring operations from the actual workload.
Use the MLOps guide for release controls, explore forecasting project planning or discuss a scoped forecasting pilot.