Why Predictive Maintenance Pilots Stall Before They Scale
The technology works in the demonstration and disappoints in the plant, and the reasons are consistent enough to plan around. What the pilot proved, what it did not, and the failure modes that only appear at fleet scale.

Predictive maintenance has an unusual reputation problem. The technology is genuinely mature — the sensors are good, the market is large and growing, and there are plants where it demonstrably saves money — and yet a large share of programmes stop after the pilot. Something works on three machines and never reaches three hundred. The reasons are consistent enough that they can be planned around, which is the only useful thing to do with them.
Start with what a pilot actually proves. A typical pilot instruments a handful of machines that somebody already suspects, collects data for a few months, and shows that the model flagged something before it broke. That is a real result and it is a weaker one than it appears. The machines were chosen because they were interesting. The data was collected during a period somebody was watching. The alerts were interpreted by the person who built the model. None of those three conditions survives contact with a fleet.
The most-cited technical failure is that anomaly detection generalises badly from clean data to real production. A study across a set of field-deployed machines found a widely-used unsupervised detector collapsing to an F1 score around 0.12 on real production signals — not a marginal degradation, a model that is wrong more often than it is right. The mechanism is not mysterious. Unsupervised detectors learn what normal looks like, and in a plant normal is not one thing: it changes with product, with ambient temperature, with tooling wear, with which operator set the machine up. A detector trained on a narrow slice of normal reports the rest of normal as anomalous, and the maintenance team learns within a fortnight to ignore it.
That leads to the second failure, which is organisational and more fatal. An alert with no owner is noise. If the system tells a planner that bearing temperature on line four is trending oddly, the useful question is what that person is supposed to do differently this week — and if the answer is "raise a work order that competes with twenty others on a schedule set by production," nothing changes and the alert quietly becomes a thing people close without reading. Programmes that scale almost always changed a process at the same time: a standing slot for condition-driven work, an agreed threshold at which a machine comes out of the schedule, someone whose job is to triage.
The third is that the economics are usually argued the wrong way round. The business case gets built on avoided catastrophic failure, which is rare, hard to attribute and impossible to prove you prevented. The value that actually accumulates is duller: fewer unnecessary scheduled interventions, longer intervals on machines that turn out to be healthy, parts ordered before the machine stops rather than during. Surveys of maintenance leaders keep finding that most plants still run predominantly reactive or calendar-based maintenance, which means the realistic comparison for a first deployment is not against a well-run predictive programme but against a calendar — and against a calendar, modest improvements are easy to find and easy to measure.
Fourth, the data foundation is almost always underestimated. A pilot can be fed by an engineer with a laptop. Three hundred machines cannot. At that point you need consistent sensor placement, an asset register that matches reality, a naming scheme that survives a plant reorganisation, and somewhere to put the data that the analytics can reach. Sites that arrive at predictive maintenance through a working historian, a unified namespace or an OPC UA information model tend to scale; sites that arrive through a pilot and then start thinking about integration tend not to.
There is also a broader caution worth taking seriously. Analysts expect a substantial share of ambitious AI-driven automation projects to be abandoned before they reach production, on grounds of unclear value and cost rather than technical failure. Predictive maintenance is not exempt, and the projects that survive that filter are usually the ones that were scoped narrowly enough to be evaluated honestly.
What all of this suggests for a first deployment. Choose the machines by consequence rather than by interest — the ones whose failure stops a line, not the ones with the most interesting vibration. Instrument a whole class of machines rather than a sample, because the value of comparison across identical assets exceeds the value of depth on one. Write down the decision the system is supposed to change, and who makes it, before writing any model. Prefer a simple, explainable indicator that a technician trusts over a sophisticated one they do not. And measure the boring number — intervention count, parts lead time, hours of unplanned stop — because that is the number that will still be defensible in a year.
The sensors have not been the limiting factor for some time. Modern MEMS vibration parts reach noise floors and bandwidths that were laboratory equipment a decade ago, and several now do processing on the die. The constraint is what happens to the output, and that is not a technology problem.