Retail 6 weeks to first deployable, 14 weeks to chain-wide Predictive AnalyticsMLOps & Deployment

Lifting forecast accuracy 22% across 1,100 stores

How a national specialty retailer replaced a spreadsheet-and-judgement forecast with a hierarchical probabilistic model that allocators actually use.

Client: National Retailer (name withheld under NDA)

Results

  • +22% Forecast accuracy (WAPE improvement)
  • −31% Stockout incidents on A-class SKUs
  • −14% Inventory carrying cost
  • 38k × 1.1k SKU-store series forecast nightly
  • 52 min Nightly full-chain run time
  • 4.1x First-year ROI

Overview

A national specialty retailer runs 1,100 stores and roughly 38,000 active SKUs. Replenishment was driven by a trailing twelve-week moving average, computed in the merchandising data warehouse, and then overridden by regional allocators using judgement and a spreadsheet. The allocators were good at their jobs; the forecast underneath them was not good at its job, and they spent most of their week compensating for it.

The symptoms were the familiar pair that always appear together: persistent stockouts on fast movers, and simultaneous overstock on seasonal and promotional items that had to be marked down.

The problem

The existing forecast had no concept of a promotion. Promotional weeks were simply absorbed into the moving average, which meant a promotion both under-forecast its own week and then over-forecast the three weeks after it. Allocators knew this and corrected manually, which means the real forecasting system was 23 people’s institutional memory — unmeasurable, unscalable, and about to retire in part.

Hierarchy consistency was broken. Store-level, region-level and chain-level forecasts were produced by different processes and did not sum. Finance planned against the chain number, supply chain executed against the store numbers, and the gap between them was a recurring quarterly argument.

The long tail dominated the SKU count. About 60% of SKU-store series sold fewer than two units a week — intermittent demand where a point forecast is nearly meaningless and a mean estimate is actively misleading.

And a previous attempt had failed. A vendor-delivered forecasting module had been implemented 18 months earlier and was essentially unused: it produced a single number per SKU-store with no explanation, no confidence range and no way for an allocator to see why it disagreed with them. That history mattered more than any technical constraint — the organisation’s bar was not “is it accurate” but “will anyone trust it this time.”

Approach

Week 1 — Discover. We measured the current forecast properly, which had never been done: WAPE at each hierarchy level, segmented by velocity class and promotional status. Baseline chain-level WAPE was 18.4%; store-SKU level was 61%. Critically, we also measured the post-override accuracy — what allocators actually ordered against. Their overrides improved accuracy on promotional items by a meaningful margin and degraded it slightly on stable items. That finding set the design brief: keep the allocators’ promotional judgement in the loop, take the stable items off their plate.

Week 2 — Architect. Three decisions. First, forecast quantiles rather than means — intermittent demand needs a distribution, and allocators need to reason about downside risk, not a single number. Second, forecast bottom-up at SKU-store and reconcile up the hierarchy with MinT, so all levels sum by construction. Third, treat promotions as explicit features with a calendar owned by merchandising, not as noise.

Weeks 3–4 — Build. LightGBM with quantile objectives, trained per store-cluster rather than per store — 1,100 individual models overfit the long tail badly in our week-two benchmarks, while a single global model lost store-level seasonality. Clustering stores by demand profile into 14 groups was the compromise that won on held-out data. Features: trailing demand at multiple windows, promotional calendar and depth, price and relative price, holiday proximity, local weather, store-local event feeds, and new-item attribute similarity for cold-start SKUs. Great Expectations gates every input table; a nightly run with silently bad POS data is worse than no run.

Week 5 — Validate. Backtesting across eight quarterly windows, including the prior year’s disrupted Q4, with accuracy segmented by velocity class, promotional status and hierarchy level. We ran the model in parallel against live allocator decisions for three weeks and — this was the step that actually sold the system internally — we showed each allocator their own historical overrides scored against what the model would have produced. Where they had been right, the model agreed. Where they had been wrong, it was usually on stable items they had not had time to think about.

Week 6 onward — Deploy and scale. Rolled out to three regions, then chain-wide over eight weeks. The allocator UI shows the forecast with its 10th/50th/90th percentiles, the top contributing features, last year’s actual for the same week, and a one-click override with a required reason code. Those reason codes became a feedback dataset and now drive the quarterly retraining.

Solution

Nightly, the system forecasts 42 million SKU-store-week series across a 13-week horizon, reconciles them to region and chain level, and publishes to the replenishment engine and the allocator UI by 5:00am local. Full chain run time is 52 minutes on Ray.

What made it stick:

Quantiles, not points. Replenishment policy consumes the service-level quantile appropriate to each velocity class. For the intermittent tail, that is the only statistically honest approach.

Reconciliation by construction. Finance and supply chain now plan against numbers that sum. The quarterly argument is gone.

Explanation per forecast. Top feature contributions, prior-year comparison and the override history are visible in the UI. Allocators override roughly 4% of forecasts now, down from effectively 100%, and every override is captured with a reason.

Exception-first workflow. Allocators no longer review a forecast list. They review flagged exceptions — unusual distributions, cold-start items, forecast-versus-constraint conflicts. Their week changed shape.

Results

Measured over the first three quarters chain-wide, against the week-one instrumented baseline:

  • +22% forecast accuracy, chain-level WAPE from 18.4% to 14.3%; store-SKU WAPE from 61% to 47%.
  • −31% stockout incidents on A-class SKUs.
  • −14% inventory carrying cost with no service-level degradation.
  • −26% markdown exposure on seasonal and promotional lines.
  • Allocator override rate down to 4% from effectively universal manual adjustment.
  • 52-minute nightly run for the full chain, well inside the 5:00am publication window.
  • 4.1x first-year ROI, dominated by the carrying-cost and markdown terms.

The retailer’s supply chain team now owns and retrains the system; KODA’s involvement ended with handover and a two-quarter advisory retainer.

Technology stack

See the stack table above. Training and inference run in the client’s own cloud account. All code, dbt models, feature definitions, trained models, backtesting harness and Terraform were delivered into the client’s repositories at close.

Let's scope your AI opportunity

A 45-minute conversation is usually enough to tell whether there is a system worth building — and what it would take. No obligation, no deck.