Shipment and order forecasting for one plant

Analysis and recommendation Jul to Aug 2026
  • Python
  • statsmodels
  • XGBoost
  • Chronos-2
  • SHAP

The plant forecast its monthly shipments with a vendor AutoML tool. I audited that tool's models first, then built a simpler forecast and evaluated it in a 54-month rolling-origin backtest covering 61 products and 16,001 scored forecasts.

39.2%WAPE over a 54-month backtest
Weighted blend39.2%
Blend + XGBoost41.6%
Best standalone ML42.0%
ML on raw volumes50.0%

Weighted absolute percentage error (WAPE) across the backtest. Lower is better.

Method and results

  1. The audit came first. Five of the six existing models had been trained on a quantity other than shipped units, and their test sets reused validation data, so there was no fair benchmark to beat.
  2. My forecast blends the order book, a 3-month average, a 12-month level and the sales plan. The weights are non-negative, sum to one and are refitted each year for every segment and horizon, using past data only.
  3. Tree models underperformed because they cannot extrapolate a level, and product volumes here span four orders of magnitude.
  4. As a leakage test, I multiplied every order line after the forecast date by 7.3 and re-ran the pipeline. The forecast changed by 0.000.
  5. For incoming orders my model reduced error from 60.7% to 31.7%. I then found that my comparison had been unfair to the plant's own formula. With a 12-month lookback instead of 6, that formula improved from 29.13% to 27.09%, so I recommended it over mine.
  6. A better order forecast would move the shipment forecast by only 0.06 percentage points, so I advised against further investment for that purpose.