Sep02
The pilot presentation is a success.
The model produces a better forecast. The dashboard looks polished. The project team demonstrates recommendations that planners could not have generated as quickly on their own.
Approval follows. So does talk of rolling it out across the network.
Six months later, the model is still running, but the planners are back in spreadsheets. Recommendations arrive too late to affect the decision. One site trusts the output; another ignores it. When the data deteriorate, nobody owns the repair. When a recommendation is wrong, nobody is certain who owns the consequence.
The technology worked.
The operating model did not.
That distinction is the subject of my new executive brief, Why supply-chain AI pilots fail to deliver lasting value.
The report draws on a structured analysis of 43 interviews from my Resilient Supply Chain podcast. Those conversations bring together people designing, selling, implementing and working with supply-chain technology. Rather than relying on one survey or a handful of headline case studies, I compared their accounts to identify recurring operating problems, successful deployments, contradictions and areas where the evidence remains weak.
The central conclusion is simple:
The unit of scale is not the model. It is the operating capability surrounding it.
A proof of concept can establish whether a model is capable of producing a useful forecast, recommendation, classification or alert.
That is valuable. It is also incomplete.
Pilots operate under unusually favourable conditions. The scope is controlled. Data are selected and prepared. A small team pays close attention. Exceptions can be handled manually. Project sponsors are available to settle disagreements. If an output looks wrong, somebody investigates.
Production removes that protective layer.
Data continue to arrive from ERP and planning systems, spreadsheets, equipment, suppliers and logistics partners. Users have competing priorities. Local conditions differ. Decisions are time-sensitive. Experienced people change roles. Integrations fail. Models encounter events poorly represented, or entirely absent, in their historical data.
At that point, model performance is only one part of the investment case.
The harder question is whether the organisation can turn the output into a better operational decision, consistently enough to justify the full cost of doing so.
That full cost includes much more than the model or software licence. It includes integration, sensing, data maintenance, governance, support, process redesign, training, exception handling and provider continuity.
A technically successful pilot can therefore conceal a commercially weak production proposition.
Across the interviews, four mechanisms received strong support: production data, connection to execution, decision authority and work design.
They should not be treated as four independent items on an AI-readiness checklist. They form a chain.
Break any critical link and value can disappear.
A pilot can be built around a curated dataset. Production depends on information that continues to arrive at the required quality, frequency and speed.
That changes the nature of the data problem.
It is no longer a one-off cleansing exercise. It is an operating responsibility with named owners, maintenance rules, latency requirements, external dependencies and recovery procedures.
In one interview, Andy Kohm of SCIP explains how contradictory and poorly maintained source data can cause AI to accelerate the wrong decisions.
The danger is not simply that the model stops working. It may continue to produce confident, plausible outputs while the quality of the underlying signal declines.
That is harder to detect, and potentially more expensive.
Even an accurate recommendation has no economic value until it changes a real decision.
This sounds obvious. In practice, it is where many analytical systems become detached from operations.
Spencer Penn of Lightsource.ai describes sourcing decisions that may be completed through email and spreadsheets before the ERP record is updated. Intelligence connected only to the formal system of record can arrive after the consequential choice has already been made.
The model has an insight.
The workflow has moved on.
Similar failures occur when an alert reaches somebody who cannot act, when a recommendation is not integrated with the system of execution or when a standard design collides with the realities of a particular warehouse, factory or transport operation.
The production test must therefore include the entire action path: where the output appears, who receives it, which system executes it, how quickly action must follow and how the result returns as feedback.
The next hand-off is organisational.
What may the AI observe? What may it recommend? What may it execute? When must a person intervene? Who may override the system? Who owns the result?
These are operating-model decisions, not technical settings.
Simon Bezrukov of Bristlecone distinguishes between automating the administration around a decision and owning its consequences. A system might gather missing information, open a ticket, propose a revised plan or select from an approved set of responses. That does not mean it should receive unrestricted authority over every decision.
“Human in the loop” does not resolve the issue on its own.
A nominal approval step can add delay without adding judgement. Equally, removing people too quickly can leave nobody capable of recognising a change in conditions, challenging a plausible error or taking responsibility for an exception.
Human involvement should reflect consequence, reversibility, confidence and the maturity of the deployment.
Authority can expand as evidence accumulates.
When people reject a new system, the explanation is often “resistance to change”.
Sometimes it is.
But that phrase can also conceal a design failure.
The software buyer may not be the person expected to use the output. A standard process may ignore legitimate site differences. An alert may add work without removing an existing task. Automation may eliminate the routine activities through which less experienced employees acquire judgement.
Training cannot compensate for a poorly designed job.
A production deployment must be developed with the people who will use and supervise it. It must be tested under realistic pressure, not only in a demonstration environment. That means rehearsing exceptions, degraded data, integration failures, provider outages and the moments when a person must take control.
The objective is not to persuade users to accept the technology.
It is to design a better way of working.
A failure-only account would be just as misleading as uncritical enthusiasm.
The interviews also contain examples of AI and automation producing operational value. These cases challenge two common assumptions.
The first is that an organisation must perfect its entire data estate before it can begin. It does not.
Focused modelling, simulation and an 80/20 approach can support a bounded decision without cleansing every field in every system.
The second is that every AI-supported action requires permanent human approval. It does not.
Opening a service ticket is not the same as committing inventory. Retrieving missing information is not the same as selecting a supplier. Adjusting a low-risk parameter within an approved range is not the same as controlling a safety-critical process.
The right degree of autonomy depends on the decision.
Some contributors also report material operational outcomes. Penn, for example, describes an unnamed automotive customer whose platform-supported categories completed sourcing 25% faster and experienced 37% less post-award cost creep than categories managed outside the platform.
That is a provider-reported comparison, not a controlled or independently audited study. It should not be treated as a universal benchmark.
It does, however, illustrate the architecture of a scaling case: a defined decision, operational integration and measures tied to the outcome.
The counterexamples do not weaken the four findings. They make them more precise.
The lesson is not to wait for perfect conditions.
It is to choose a consequential but bounded decision, test the complete operating chain and expand only when the evidence supports expansion.
Before approving another AI pilot, leaders should be able to answer eight questions:
Together, those answers form a bounded production hypothesis.
That is a better basis for investment than a loosely defined technology experiment. It forces the team to test whether the organisation can operate the capability, not merely whether the model can generate an impressive output.
It also creates a legitimate way for a pilot to succeed without being scaled unchanged.
A well-designed experiment may show that the expected value is not there. It may reveal an integration cost that changes the economics. It may demonstrate that process redesign would solve more of the problem than machine learning.
That is useful knowledge.
A pilot fails when it neither improves a material outcome nor resolves an uncertainty worth funding.
I created the report by returning to recent Resilient Supply Chain interviews and retesting an earlier set of conclusions rather than assuming they were correct.
The analysis examined 43 interviews and 154 relevant transcript sections published between 3 November 2025 and 31 August 2026. It deliberately searched for evidence that challenged the emerging argument: successful deployments, alternative explanations, contradictory experiences and differences between what technology providers promise and what organisations report in practice.
This matters because retrieval is not proof. Finding several people making similar claims does not establish how common a problem is across the industry. Nor does a persuasive customer story become an independent benchmark simply because it includes an impressive number.
The strength of the research lies elsewhere.
It brings together detailed accounts from people working across supply-chain technology and operations, compares the mechanisms they describe and tests whether the initial explanation survives contact with exceptions and disagreement.
Four connected mechanisms emerged with strong support: production data, connection to execution, designed authority and redesigned work.
A fifth finding, value discipline, is suggestive rather than equally established. Several contributors argue, persuasively, that AI initiatives should begin with a valuable business problem and measurable baseline. But the interviews do not contain enough retrospective evidence from cancelled pilots to establish weak ROI discipline as a primary cause as strongly as the other four.
The research is qualitative. The interview participants were not randomly selected, and technology providers and advisers are more heavily represented than operators. The findings cannot tell us what proportion of all supply-chain AI pilots fail for a particular reason.
They can do something more useful for an executive deciding whether to fund the next stage: identify recurring failure mechanisms, reveal important exceptions and sharpen the questions that should be answered before more capital and organisational attention are committed.
Return to that successful pilot presentation.
The forecast may genuinely be better. The recommendation may be useful. The underlying technology may be excellent.
Lasting value will still depend on what happens next: maintaining the inputs, connecting the output to action, allocating authority, redesigning the work, managing exceptions and measuring whether performance improves after the project team steps away.
So the next time a pilot team presents an impressive result, do not ask only:
“Did the model work?”
Ask:
“Can the organisation operate it?”
That is the production question.
And it is the question that separates an interesting demonstration from a capability worth scaling.
Read the research and download the free executive brief.
No registration required.
By Tom Raftery
Keywords: AI Orchestration, Digital Transformation, Supply Chain
Your Supply-Chain AI Pilot Worked. So Why Didn’t It Scale?
Why Every AI Program Needs a Project Manager
The Line Between Booking Freight and Controlling It
Why do I need an AI strategy to achieve AI adoption?
How Fresh is Your Leadership