OTIF — on time, in full — is the only service metric most customers actually care about, and it is the one most often reported in a way that makes an operation look better than the customer experiences it.
This is what the number really measures, the four definitional choices that move it by ten points without anything changing on the floor, and what genuinely improves it.
The arithmetic is multiplicative, and that is the whole problem
On time and in full are two separate tests, and an order has to pass both. So they multiply.
A warehouse hitting <strong>95% on time</strong> and <strong>97% in full</strong> is not at 96%. It is at <strong>0.95 × 0.97 = 92.2% OTIF</strong>. Two respectable numbers produce a mediocre one, and the gap widens the more lines an order has.
This is why operations are often genuinely surprised by their OTIF. Each department reports its own metric honestly — transport reports on-time, the warehouse reports fill rate — and nobody multiplies. The first useful thing most companies do with OTIF is simply calculate it correctly.
Four definitions that change the number without changing reality
1. Per order or per line?
An order with ten lines where nine ship complete is 90% at line level and <strong>0% at order level</strong>, because the order was not complete. Both are defensible; they are not comparable. Retail customers almost always measure at order or even delivery level, because that is what they receive. Suppliers almost always report at line level, because it is kinder.
2. Whose date?
On time against <em>the date the customer requested</em>, <em>the date you confirmed</em>, or <em>the date you last promised after a change</em>? Measuring against your own confirmed date is the most common choice and the most flattering: any order you re-promise resets its own clock, so a chronically late line can score 100%.
If your OTIF looks excellent and your customers complain anyway, this is usually why. The honest version measures against the <strong>original requested date</strong>, and reports re-promises as a separate figure rather than absorbing them.
3. What counts as in full?
Does a substitution count? A short shipment the customer accepted? An over-delivery? Rounding to a full case when the customer ordered eaches? Each of these is a policy decision, and each moves the number. Write them down, because the customer has written down their version.
4. What is the delivery window?
Same-day is a point; most customers work with a window. A two-day window forgives a lot that a same-day target does not, and "on time" against a window that starts the day after the request is a different promise entirely.
OTIF-D and why customers impose it
Large retailers do not track OTIF to be informed; they track it to charge. Supplier scorecards commonly carry financial penalties for missing a threshold, and the threshold is measured <strong>their</strong> way — usually per delivery, against the original window, with no credit for a re-promise.
The practical consequence is that your internal OTIF and your customer's OTIF for the same shipments are different numbers, and only one of them generates a deduction. If you supply retail, the metric worth reporting internally is theirs, reconstructed as closely as you can, even if it looks worse.
What actually moves OTIF (usually not the warehouse)
When OTIF is poor, the instinct is to look at picking accuracy and dispatch cut-offs. Occasionally that is right. More often the causes sit upstream of the warehouse entirely:
- <strong>Promise-setting.</strong> Availability is confirmed against stock that is physically present but already allocated, or against a replenishment that has not arrived. The order is late the moment it is accepted.
- <strong>Inventory positioning.</strong> The stock exists, in the wrong location or the wrong unit. A full pallet in a regional site does not help an eaches order due tomorrow from another one.
- <strong>Demand variability meeting a fixed replenishment cycle.</strong> If you replenish weekly and demand moves daily, a proportion of orders will always meet an empty pick face regardless of how well anyone picks.
- <strong>Supplier reliability</strong> passed straight through. Inbound OTIF from your own suppliers sets a ceiling on your outbound OTIF that no amount of warehouse effort raises.
This is the diagnostic value of splitting the metric. On-time failures and in-full failures have almost entirely different causes: on-time is usually transport, cut-offs and promise-setting; in-full is availability, allocation and replenishment. Reporting them separately, then multiplying, tells you which department owns the problem.
The improvement that is not an improvement
The fastest way to raise OTIF is to promise later dates. Quote five days instead of three and on-time rises immediately, without a single process change.
Sometimes that is exactly right — a promise you can keep beats an optimistic one you cannot, and customers generally plan around reliability rather than speed. But it must be a decision someone takes deliberately, not a drift that happens because OTIF is on a dashboard and lead time is not. <strong>Track average promised lead time alongside OTIF</strong>, or you will optimise one by quietly degrading the other.
Where forecasting and automation genuinely help
OTIF is a measurement problem before it is a technology problem, and most of the gain in the first months comes from calculating it correctly and splitting it by cause. After that, two applications earn their keep:
- <strong>Demand forecasting feeding replenishment</strong>, so that the pick face is stocked ahead of demand rather than after a stockout. This attacks in-full directly and is the highest-value place for a model in most operations.
- <strong>Availability checking at order entry</strong> that reflects allocated stock and inbound timing rather than a raw on-hand figure, so the promise is achievable when it is made.
Neither requires a large programme. Both require order and delivery history that is complete enough to measure against, which is the same data you need to report OTIF honestly in the first place.
Inbound OTIF: the scorecard you should be running on your own suppliers
Almost every company that is measured on OTIF by its customers fails to measure its own suppliers the same way. That asymmetry is expensive, because inbound reliability sets a hard ceiling on outbound performance.
The mechanism is simple. If a supplier delivers late or short, you either hold the order — an outbound OTIF failure you will be charged for — or you expedite, which costs margin. Either way the failure originated outside your building and arrives on your scorecard.
- Measure inbound on the <strong>same definition</strong> your customers use on you. Anything softer tells you less than nothing.
- Split it by supplier and by article, not just by supplier. One unreliable line from an otherwise good supplier is a specific conversation, not a relationship problem.
- Convert it into safety stock. A supplier at 80% inbound OTIF is not a supplier problem to solve first — it is a buffer you have not sized, and knowing the number is what lets you size it deliberately instead of by anxiety.
Suppliers respond to being measured far more than to being asked. Sending a monthly figure, calculated the same way every month, changes behaviour more reliably than escalation does.
Beyond OTIF: the perfect order rate
OTIF asks two questions. The metric practitioners graduate to — <strong>perfect order rate</strong> — asks four: on time, in full, <em>undamaged</em>, and <em>with correct documentation</em>.
The last two matter more than they sound. A delivery that arrives on time and complete but with a wrong packing list generates the same customer-service work as a short shipment, and in regulated flows a documentation error can stop the goods being accepted at all. Because the tests multiply, adding two more drags the number down again: four tests at 97% each is 88.5%.
You do not need to report perfect order rate to benefit from the idea. The useful part is the recognition that <strong>OTIF is a floor, not a ceiling</strong> — an operation can hit its OTIF target and still generate steady complaint volume from the two failure types the metric does not look at.
How to report it so it changes something
A single OTIF percentage on a dashboard is close to inert: it tells you the score without telling anyone what to do.
- Report the <strong>product and both components</strong> together, always. The components carry the diagnosis.
- Attribute every failure to one cause code at the point it fails, not retrospectively. Retrospective attribution reliably drifts toward whichever department is not in the room.
- Show the <strong>customer view alongside the internal view</strong> for your largest accounts. The gap is the most actionable number in the report.
- Track average promised lead time on the same chart, so that OTIF gains bought by promising later are visible as what they are.
How we approach it
We start with a <strong>€2,500 audit</strong>: reconstruct OTIF the way your largest customer measures it, split on-time from in-full, and attribute the failures to cause. That alone frequently reallocates the improvement effort away from the warehouse. If a build follows — usually forecasting into replenishment, or an availability check at order entry — a <strong>€20,000 proof of concept</strong> runs on your history for four to six weeks, and production starts from <strong>€50,000</strong>.
Frequently asked questions
How is OTIF calculated exactly?
Multiply the on-time rate by the in-full rate for the same set of orders. They are two independent tests and an order must pass both, so 95% on time and 97% in full gives 92.2% OTIF, not 96%.
- Decide the unit first — per order line, per order or per delivery. The same shipments produce very different numbers.
- Use one consistent date basis. Mixing requested and re-promised dates makes the series meaningless over time.
- Report on-time and in-full separately as well as the product, because the two have different owners.
What is a good OTIF score?
It depends entirely on the definition, which is why cross-company comparisons are close to worthless without knowing how each one measures.
- Retail supplier scorecards commonly set thresholds in the mid-to-high nineties, measured per delivery against the original window — a harder test than most internal reporting.
- A line-level figure against re-promised dates will always look better than an order-level figure against requested dates. Same operation, different number.
- The more useful target is your own trend on a fixed definition, plus the gap between your number and your customer's for the same shipments.
Why is our OTIF good but customers still complain?
Almost always a definition gap. The usual culprit is measuring on time against your own confirmed or re-promised date rather than the date the customer asked for.
- Every re-promise resets the clock, so an order that slipped twice can still count as on time.
- Line-level measurement hides incomplete orders: nine of ten lines is 90% to you and a failed delivery to them.
- Reconstruct the metric their way for one month. The gap between the two numbers is the size of the reporting problem.
Does improving OTIF mean investing in the warehouse?
Often not. Split the metric before spending anything: on-time failures usually trace to transport, cut-offs and promise-setting, while in-full failures trace to availability, allocation and replenishment timing.
- If in-full dominates, the constraint is upstream — forecasting, inventory positioning or supplier reliability — and picking faster changes nothing.
- Inbound OTIF from your own suppliers caps your outbound figure. Measure it before assuming the problem is internal.
- The cheapest first move is almost always measuring correctly and attributing failures to cause.