Supply Chain & Logistics Management

Forecasting Methods and Sales and Operations Planning

Module 2

Almost every supply chain decision rests on an estimate of future demand, and this lesson covers the full forecasting toolkit: error metrics, moving averages, exponential smoothing, trend and seasonality models, plus aggregate planning and S&OP.

1. Role of Forecasting

Demand Planning, Forecasting and Sales and Operations Planning: module overview infographic

Forecasting answers what customers will want, when, and where. Inventory, capacity, transportation, and sourcing decisions all depend on demand estimates; supply chain planning is fundamentally about matching supply to demand.

Process TypeDefinitionForecasting Role
Push ProcessPerformed in anticipation of demand, before an order is placedBasis for manufacturing, transporting, and stocking decisions
Pull ProcessPerformed in response to an actual orderNeeded to plan capacity and inventory availability so the system responds quickly

Paint retail case: final mixing of base paint and dyes happens after the customer selects a shade (pull), but the retailer must already stock base paint and dyes (push) because factory ordering at the moment of request would take too long. The store forecasts stocking levels, the factory forecasts base paint production, and upstream suppliers forecast raw materials.

Collaborative forecasting: independent forecasting at separate stages creates misalignment, causing stockouts or excess inventory. Partners must share planning information and align on a single common demand view. Example: a beverage company's promotion must be reflected in the joint forecast, otherwise the bottler schedules for normal weeks and a shortfall follows.

HorizonTimeframeSupported DecisionsRole
Long RangeMonths, yearsCapacity planning, warehouse locations, new product launchesBuilding the system
Short RangeDaily, weeklyReplenishment, staffing, production schedules, delivery planningRunning the system

Seasonal retailers integrate weather forecasts (rain raises umbrella demand; heatwaves raise cold beverage and cooler demand). Auto dealers must stock the right model and variant mix because buyers expect immediate delivery; a wrong forecast leaves sitting inventory and tied-up working capital.

Expected demand vs. forecast error: planners track both the expected level of demand and the uncertainty around it. Forecast error drives the buffers in the system. When error is high, managers choose among safety inventory, flexible capacity, faster replenishment, or demand shaping through pricing.

Organizational alignment: sales teams are optimistic, operations conservative, finance cost-focused. Planning with mismatched functional forecasts produces excess costs or poor service, so functional consensus matters.

2. Features of Forecasts and the Forecasting Process

Four Common Features of All Forecasts

FeatureDefinitionImplication
Continuity AssumptionThe system that created the past will continue into the futureModels cannot anticipate weather shocks, competitor moves, policy changes; manual override needed
Inherent ImperfectnessRandomness prevents perfect forecastsQuantify error and maintain buffers
Aggregation AccuracyForecasts are more accurate for groups than individual itemsFluctuations cancel out over categories or regions
Horizon SensitivityAccuracy decreases as the horizon lengthensFlexible chains exploit accurate short-horizon forecasts

Elements of a Good Forecast

Timely, mathematically accurate (with stated error estimates), reliable, in meaningful units, written and shared, simple enough to be trusted, and cost-effective (benefits exceed data and system costs).

Consequences of Inaccuracy

DeviationConsequences
Under-forecasting (too low)Shortages, stockouts, missed deliveries, production disruptions, poor service
Over-forecasting (too high)Excess inventory, idle capacity, high holding costs, markdowns

Bad forecasts create amplification waves: a stage reacting with sudden order changes causes upstream partners to misinterpret them, amplifying variability. Mitigated by joint forecasting and information sharing.

Six-Step Forecasting Process

The six-step forecasting process
  1. Determine the purpose (decision supported). 2. Establish the time horizon. 3. Obtain and clean data (errors, outliers, comparability). 4. Choose a method. 5. Generate the forecast. 6. Monitor errors and update, forming a closed feedback loop.

3. Demand Components and Forecast Accuracy

Demand ComponentDefinitionExamples / Treatment
Horizontal (Level)Demand fluctuates around a stable averageMature staples, steady B2B consumables
TrendLong-term systematic upward or downward movementUpdate the baseline to prevent chronic shortages or excess
SeasonalityRegular repeating patterns with fixed, known periodicity (daily to yearly)Summer ice cream, monsoon umbrellas, weekend restaurant surges
CyclicalWave-like movements longer than a year, periodicity not fixedCapital equipment, luxury goods, construction; managed via scenario planning
Irregular VariationOne-off spikes or dropsSevere weather, strikes, pandemics; flagged and removed from baseline
Random (Noise)Residual after all patternsManaged with safety stock and capacity buffers

Forecast Error Metrics

ƒForecast error
Et=AtFtE_t = A_t - F_t
Where: A_t is actual demand and F_t is the forecast for period t. Positive error means under-forecasting; negative error means over-forecasting.
MetricFormulaInterpretation
Mean Error (Bias)ME=1nEtME = \frac{1}{n}\sum E_tSystematic over- or under-forecasting
MAD$MAD = \frac{1}{n}\sumE_t
MSEMSE=1nEt2MSE = \frac{1}{n}\sum E_t^2Penalizes large errors more heavily
MAPE$MAPE = \frac{1}{n}\sum \frac{E_t
ƒMean error (bias)
ME=1nEt\mathrm{ME} = \frac{1}{n}\sum E_t
ƒMean absolute deviation (MAD)
MAD=1nEt\mathrm{MAD} = \frac{1}{n}\sum |E_t|
ƒMean squared error (MSE)
MSE=1nEt2\mathrm{MSE} = \frac{1}{n}\sum E_t^2
ƒMean absolute percentage error (MAPE)
MAPE=1nEtAt×100\mathrm{MAPE} = \frac{1}{n}\sum \frac{|E_t|}{A_t} \times 100
Where: n is the number of periods compared and E_t is the forecast error in period t.

Worked example. Given, over n = 8 days for a quick-commerce SKU: absolute error sum 22, squared error sum 76, absolute percentage error sum 10.26.

MAD=228=2.75 units/day\mathrm{MAD} = \frac{22}{8} = 2.75 \text{ units/day}
MSE=768=9.5\mathrm{MSE} = \frac{76}{8} = 9.5
MAPE=10.268=1.28%\mathrm{MAPE} = \frac{10.26}{8} = 1.28\%

Answer: MAD = 2.75 units/day, MSE = 9.5, MAPE = 1.28%, a highly accurate baseline.

4. Forecasting Approaches

CategoryBasisUse Cases
QualitativeHuman judgment, expert opinion, surveysNew launches, limited history, long-term strategic shifts
QuantitativeHistorical data, time series, causal regressionDay-to-day operations, stable data-rich environments

Firms often combine: a quantitative baseline plus qualitative managerial adjustments for promotions or competitor moves.

Qualitative Methods

MethodSourcingStrengthsRisks
Executive OpinionSenior manager groupCross-functional strategic insightStrong personalities or hierarchy bias dominate
Sales Force OpinionsSales team customer contactEarly market signalsConfusing interest with buying; quota incentives distort
Consumer SurveysDirect customer queriesDirect demand insight for new productsCostly, slow, non-response bias, stated vs. actual gap
Delphi MethodAnonymous iterative expert questionnairesLimits loud voices, objective consensus for long-range questionsTime-consuming, depends on expert selection

Simple and Weighted Moving Averages

A time series is a sequence of observations at regular intervals; models assume the near future behaves like the recent past.

ƒSimple moving average
Ft=MAn=At1+At2++AtnnF_t = \mathrm{MA}_n = \frac{A_{t-1} + A_{t-2} + \dots + A_{t-n}}{n}
Where: n is the window length and A_{t-1} to A_{t-n} are the most recent n actual demands.

Larger n smooths fluctuations but reacts slowly (systematic lag); smaller n reacts quickly but chases noise.

Worked example. Given: demand over the last three weeks was 43, 40, 41.

F6=43+40+413=41.33 unitsF_6 = \frac{43 + 40 + 41}{3} = 41.33 \text{ units}

If actual week 6 demand is 38:

F7=40+41+383=39.67 unitsF_7 = \frac{40 + 41 + 38}{3} = 39.67 \text{ units}

Answer: F6 = 41.33 units; F7 = 39.67 units.

Simple moving averages weight all observations equally, so the oldest value influences as much as the newest, creating lag. The weighted moving average assigns unequal weights (summing to 1), usually heaviest on recent data. SMA is the special case with equal weights.

Worked example. Given: weights 0.4, 0.3, 0.2, 0.1 on demands 41, 40, 43, 40.

F6=0.4×41+0.3×40+0.2×43+0.1×40=16.4+12+8.6+4=41 unitsF_6 = 0.4 \times 41 + 0.3 \times 40 + 0.2 \times 43 + 0.1 \times 40 = 16.4 + 12 + 8.6 + 4 = 41 \text{ units}

Answer: F6 = 41 units. More sensitive to recent change, but weight selection is subjective trial and error.

5. Exponential Smoothing

Each new forecast adjusts the previous forecast by a fraction of the previous forecast error.

ƒSimple exponential smoothing (error-correction form)
Ft=Ft1+α×(At1Ft1)F_t = F_{t-1} + \alpha \times (A_{t-1} - F_{t-1})
Where: α is the smoothing constant between 0 and 1, F_{t-1} is the previous forecast and A_{t-1} the previous actual demand.
ƒSimple exponential smoothing (weighted-average form)
Ft=α×At1+(1α)×Ft1F_t = \alpha \times A_{t-1} + (1 - \alpha) \times F_{t-1}

The new forecast is a weighted average of the latest actual (weight α) and the previous forecast (weight 1 - α). Because the previous forecast embeds all prior demand, weights decay exponentially over history.

Alpha ValueSmoothingResponsivenessBest Use
Close to 0 (0.05-0.10)High smoothing, ignores short-term errorsSlow reactionStable SKUs where noise dominates
Close to 1 (0.40-0.50)Low smoothing, adjusts heavily to last errorJumpy, chases noiseVolatile SKUs with shifting demand averages

Worked example. Given: previous forecast 42 units, actual demand 40, α = 0.10.

F=42+0.1×(4042)=41.8 unitsF = 42 + 0.1 \times (40 - 42) = 41.8 \text{ units}

If the next actual is 43:

F=41.8+0.1×(4341.8)=41.92 unitsF = 41.8 + 0.1 \times (43 - 41.8) = 41.92 \text{ units}

Answer: the new forecasts are 41.8 units, then 41.92 units. Each forecast steps in the direction of the last error. With α = 0.1 the update is "90% stick with the old forecast, 10% listen to the latest actual":

Illustration
Ft=0.9×Ft1+0.1×At1F_t = 0.9 \times F_{t-1} + 0.1 \times A_{t-1}

Firms pick α via judgment or historical trial and error minimizing MAD, MSE, or MAPE (common values run about 0.05 to 0.5). Initialization options for the starting forecast: naive (F2 = A1), average of first few actuals, or a managerial estimate.

6. Comparing Methods

Worked example (setup). Given actuals: A1 = 42, A2 = 40, A3 = 43, A4 = 40.

Calculation (MA2)
F3=42+402=41  (error +2)F_3 = \frac{42 + 40}{2} = 41 \; (\text{error } +2)
F4=43+402=41.5  (error 1.5)F_4 = \frac{43 + 40}{2} = 41.5 \; (\text{error } -1.5)
Calculation (WMA2, weights 0.6 recent and 0.4 prior)
F3=0.6×40+0.4×42=40.8  (error +2.2)F_3 = 0.6 \times 40 + 0.4 \times 42 = 40.8 \; (\text{error } +2.2)
F4=0.6×43+0.4×40=41.8  (error 1.8)F_4 = 0.6 \times 43 + 0.4 \times 40 = 41.8 \; (\text{error } -1.8)
Calculation (single smoothing, α = 0.10, F2 = A1 = 42)
F3=0.9×42+0.1×40=41.8F_3 = 0.9 \times 42 + 0.1 \times 40 = 41.8

Errors from each method then feed MAD, MSE, and MAPE for the head-to-head comparison.

All models must be evaluated over the exact same historical period, excluding warm-up periods needed by moving averages (e.g., comparing MA2 and single smoothing from period 3 through 11). The best method depends on the criterion: in the comparative study, WMA2 achieved lowest MAD and lowest MAPE, while single exponential smoothing (α = 0.10) achieved lowest MSE, preferred when large misses cause severe crises.

7. Techniques for Trend

When a trend is present, flat methods lag: under an upward trend they systematically under-forecast (stockouts, expedited shipping); under a downward trend they over-forecast (excess inventory, markdowns).

ƒLinear trend forecast
Ft=a+b×tF_t = a + b \times t
Where: a is the intercept at t = 0, b is the slope or average change per period, and t is the period index. Estimated by least squares regression (Excel Data Analysis - Regression, demand as Y, time index as X).

Worked example. Given: regression on 10 weeks of unit sales gives a = 699.4, b = 7.5 (sales grow 7.5 units/week).

F11=699.4+7.5×11=781.9F_{11} = 699.4 + 7.5 \times 11 = 781.9
F12=699.4+7.5×12=789.4F_{12} = 699.4 + 7.5 \times 12 = 789.4

Answer: F11 = 781.9 (782 units); F12 = 789.4 (790 units).

Practical cautions: avoid extrapolating far outside the observed range, verify linearity holds within the horizon, and confirm residuals look random (patterns indicate missing model structure).

8. Trend-Adjusted Smoothing (Holt's Method)

Holt's Method (double exponential smoothing) separately updates a level S_t (current baseline) and a trend T_t (movement per period), preventing the systematic lag of single smoothing.

ƒTrend-adjusted forecast
TAFt+1=St+Tt\mathrm{TAF}_{t+1} = S_t + T_t
Where: S_t is the smoothed level at the end of period t and T_t is the smoothed trend per period.
ƒLevel update
St=TAFt+α×(AtTAFt)S_t = \mathrm{TAF}_t + \alpha \times (A_t - \mathrm{TAF}_t)
ƒTrend update
Tt=Tt1+β×(TAFtTAFt1Tt1)T_t = T_{t-1} + \beta \times (\mathrm{TAF}_t - \mathrm{TAF}_{t-1} - T_{t-1})

α smooths the level; β smooths the trend. High values react fast; low values stay stable.

Worked example. Given: A1..A4 = 700, 724, 720, 728, with α = 0.4 and β = 0.3.

Initialization: net change is 28 over 3 steps, so

T4=283=9.33T_4 = \frac{28}{3} = 9.33
S4=A4=728S_4 = A_4 = 728
TAF5=728+9.33=737.33\mathrm{TAF}_5 = 728 + 9.33 = 737.33

With A5 = 740:

S5=737.33+0.4×(740737.33)=738.40S_5 = 737.33 + 0.4 \times (740 - 737.33) = 738.40
T5=9.33+0.3×(737.337289.33)=9.33T_5 = 9.33 + 0.3 \times (737.33 - 728 - 9.33) = 9.33
TAF6=738.40+9.33=747.73 units\mathrm{TAF}_6 = 738.40 + 9.33 = 747.73 \text{ units}

Answer: TAF6 = 747.73 units.

9. Seasonality

Model TypeStructureDefinitionBest Use
AdditiveDemand = Trend + SeasonalityConstant absolute addition or subtraction (+20 units, -10 units)Seasonal swings that stay constant as business scales
MultiplicativeDemand = Trend × SeasonalityMultiplier on baseline (1.2×, 0.75×)Swings that scale proportionally with volume
ƒAdditive seasonality model
Demand=Trend+Seasonality\text{Demand} = \text{Trend} + \text{Seasonality}
Where: the seasonal component is a constant absolute amount added to or subtracted from the trend, suited to swings that stay the same size as the business scales.
ƒMultiplicative seasonality model
Demand=Trend×Seasonality\text{Demand} = \text{Trend} \times \text{Seasonality}
Where: the seasonal component is a multiplier on the baseline, suited to swings that scale proportionally with volume.

Multiplicative multipliers are seasonal relatives (indices): 1.2 means 20% above baseline average; 0.75 means 25% below.

Two workflows: deseasonalize (divide actual demand by the seasonal relative to reveal the clean trend) and forecast with seasonality (forecast the baseline, then multiply by the seasonal relative).

Worked example. Given seasonal relatives: Q1 = 1.2, Q2 = 1.1, Q3 = 0.75, Q4 = 0.95.

Deseasonalizing period 1 (Q1, actual 158.4) and period 3 (Q3, actual 110):

158.41.2=132 gallons\frac{158.4}{1.2} = 132 \text{ gallons}
1100.75=146.7 gallons\frac{110}{0.75} = 146.7 \text{ gallons}

Forecasting with seasonality, using the baseline trend:

Ft=124+7.5tF9=191.5,  F10=199F_t = 124 + 7.5t \quad \Rightarrow \quad F_9 = 191.5, \; F_{10} = 199

Applying relatives (period 9 is Q1, period 10 is Q2):

Final9=191.5×1.2=229.8 gallonsFinal_9 = 191.5 \times 1.2 = 229.8 \text{ gallons}
Final10=199×1.1=218.9 gallonsFinal_{10} = 199 \times 1.1 = 218.9 \text{ gallons}

Answer: final forecasts are 229.8 gallons (period 9) and 218.9 gallons (period 10).

Worked example (computing relatives). Step 1: season totals Q1 = 60, Q2 = 30, Q3 = 66, Q4 = 84. Step 2: season averages 20, 10, 22, 28. Step 3: overall average:

20+10+22+284=20\frac{20 + 10 + 22 + 28}{4} = 20

Step 4: relatives Q1 = 1.0, Q2 = 0.5, Q3 = 1.1, Q4 = 1.4.

Answer: Q4 is the peak at 40% above average; Q2 runs 50% below.

10. Associative Techniques

Associative forecasting models demand as a function of observable driver variables rather than time alone:

ƒSimple linear regression (associative model)
Y=a+b×XY = a + b \times X

Where: Y is the predicted demand, X is the driver variable, and a and b are the intercept and slope estimated by regression. Multiple simultaneous predictors (price, promotional intensity, weather) require multiple linear regression, larger datasets, and care against overfitting; saturation or threshold effects require nonlinear models or transformations.

11. Aggregate Planning

Aggregate planning is an intermediate-horizon process (typically 3 to 18 months) grouping products into families or total volumes rather than SKUs, translating forecasts into capacities, production levels, and inventory strategies ahead of long-lead-time constraints.

LeverDefinitionTrade-off
Production RateVolume planned per periodRamping needs workforce or overtime; flat rates need inventory or backlogs
Workforce LevelLabor force sizingHiring has training lead times; downsizing carries layoff and morale costs
OvertimeCapacity beyond regular hoursShort-term flexibility at an overtime wage premium
SubcontractingExternal production capacityInstant expansion at a price premium, subject to supplier constraints
Planned InventoryCarrying stock from low to high demand periodsStable production but ties up capital and adds holding and obsolescence costs
Backlog and StockoutsDelayed or lost unmet demandMinimizes inventory costs but damages service, margins, future demand

Three cost domains are balanced: capacity costs (wages, overtime premiums, hiring, training, layoffs, subcontracting), inventory costs (holding, storage, tied-up working capital, shrinkage, obsolescence), and backlog/stockout costs (wait penalties, lost margins, goodwill erosion).

12. Sales and Operations Planning (S&OP)

S&OP is the cross-functional process aligning the demand plan, supply plan, and financial plan into a single operational plan across marketing, finance, procurement, and operations.

Without S&OP, functions optimize in isolation: marketing wants holiday promotion spikes, operations wants stable rates and minimal overtime, distribution worries about warehouse throughput and trucks, procurement targets bulk buys against packaging lead times, and finance minimizes working capital. S&OP resolves these trade-offs by testing which promotions are operationally feasible and financially viable.

StrategyCore LeverCharacteristicsBest FitRisksExamples
ChaseCapacityProduction synchronized with demand via hiring/firing, shifts, subcontractingCheap capacity changes, expensive inventory, short lifecyclesLabor regulations, training lead times, moraleServices, call centers, gig delivery
FlexibilityUtilizationStable workforce; capacity varied via overtime and flexible schedulingAvoids hiring/layoff churn with output flexibilityOvertime premiums, hour limitsManufacturing with clear overtime guidelines
LevelInventoryConstant production and headcount; inventory or backlogs absorb swingsStable environments with cheap inventoryTied-up capital, storage limits, obsolescenceCommodities, capital-intensive manufacturing

In practice firms use hybrid strategies (some overtime, some inventory, some subcontracting) because the cheapest lever varies by context.

Memory hook: The three S&OP strategies by lever: "CFL bulb" - Chase changes Capacity, Flexibility flexes utilization (overtime), Level leans on inventory.

13. Exam Essentials

  • Expected demand vs. forecast error: expected demand is the baseline volume; forecast error is the uncertainty and directly sizes safety buffers.
  • SMA vs. WMA: SMA weights all window observations equally (lag); WMA emphasizes recent data (responsive but arbitrary weights).
  • Single smoothing vs. Holt's: single tracks one level and lags during trends; Holt's recursively updates level (S_t) and trend (T_t).
  • Additive vs. multiplicative seasonality: constant absolute change vs. proportional scaling factor.
  • SKU planning vs. aggregate planning: individual items day-to-day vs. families/volumes over 3-18 months.
  • Chase / Flexibility / Level: vary capacity via headcount, vary utilization via overtime, or keep production constant and buffer with inventory.
  • Key terms: collaborative forecasting, Delphi method, smoothing constants α (level) and β (trend), seasonal relatives, deseasonalizing, associative forecasting, S&OP.