Predictive maintenance means monitoring a machine's condition (vibration, temperature, motor current) with sensors, spotting the drift towards failure early and scheduling maintenance when it is actually needed. Reactive maintenance acts after a breakdown and preventive maintenance follows a calendar; predictive maintenance decides on the machine's own data, cutting unplanned downtime, unnecessary part changes and emergency labour.
What really happens in the maintenance workshop
In most plants, maintenance swings between two extremes. When a critical pump or compressor fails, the line stops, the maintenance team switches into emergency mode, a supplier is chased if the spare part is not on the shelf, and overtime is booked. Afterwards the reflex is to tighten the preventive schedule: bearings are swapped at fixed hours even when they are healthy, and machines that are running well are stopped for service.
Preventive maintenance is progress, but it has two blind spots. It misses failures that begin before the scheduled date, and it spends money and production time replacing parts with life left in them. Maintenance managers often cannot answer a simple question: "Is this machine in better or worse shape today than it was last month?"
Predictive maintenance answers that question. A baseline of normal operation is established, deviations from it are tracked, and maintenance decisions follow the machine's actual condition rather than the calendar. In our Industry 4.0 roadmap for SMEs, this is the prediction stage that follows data collection, visibility, production management and quality.
The cost of leaving it unsolved
Unplanned downtime is the most expensive outcome of a maintenance strategy.
According to Siemens' The True Cost of Downtime 2024, the world's 500 largest companies lose almost $1.4 trillion a year to unplanned downtime, equivalent to 11% of their revenues.
The same report finds that large plants average 25 unplanned downtime incidents a month and lose 27 hours of production a month. The good news is that both figures have fallen since 2019: nine out of ten respondents now do some form of condition monitoring, and almost half have dedicated predictive maintenance teams.
The most widely cited official source on returns is the US Department of Energy's maintenance guide. The U.S. Department of Energy Operations & Maintenance Best Practices Guide, Chapter 5 (2010) states that a properly functioning predictive maintenance programme can save 8–12% over preventive maintenance alone, and that facilities relying heavily on reactive maintenance could see savings opportunities above 30–40%.
Independent surveys cited in the guide report industry-average reductions of 35–45% in downtime and 25–30% in maintenance costs. These are averages, not promises for every site, but the direction is clear.
Maintenance strategies compared
| Strategy | What triggers the work? | Strength | Weakness |
|---|---|---|---|
| Reactive (run to failure) | The machine breaks down | Needs no planning | Unplanned downtime, emergency labour, secondary damage |
| Preventive / scheduled | Calendar or running hours | Simple and predictable | Misses early failures, replaces healthy parts |
| Condition-based | A measured value crosses a threshold | Looks at real condition | Poorly set thresholds alarm too late or too often |
| Predictive | Data trends and model forecasts | Sees failures developing, makes maintenance plannable | Requires sensors, data infrastructure and analytical discipline |
Predictive maintenance is not needed for every machine. For cheap, redundant equipment whose failure does not affect output, run-to-failure can be a deliberate and sensible choice. The real value lies in critical equipment whose failure is expensive and stops the line.
How predictive maintenance works
The process runs in four steps:
- Sensor installation: vibration, temperature, current and, where relevant, oil-quality sensors are fitted to critical equipment such as motors, pumps, compressors, bearings and gearboxes. Wired or wireless connections are both options; for remote or scattered equipment we compare the choices in LoRa vs NB-IoT vs cellular.
- Continuous data collection: sensors generate data at a defined sampling rate; an edge gateway on site pre-processes it and forwards it to the server. Why that pre-processing happens on site is explained in edge computing in manufacturing. If the internet connection drops, a local buffer takes over.
- Analysis: first, a baseline of normal operation is established. An AI model then evaluates changes in the vibration spectrum, anomalies at bearing frequencies and temperature rises together. Technical detail such as which fault shows at which frequency, ISO 20816 limits and sensor choice is covered in vibration analysis for fault detection. The approach that combines sensor data with a physical model of the machine to test scenarios is discussed in our digital twin article.
- Alerts and action: when a threshold is crossed, the maintenance team receives an email, SMS or app notification; with integration into a computerised maintenance management system (CMMS), a work order opens automatically. How a predicted failure becomes a work order assigned to a technician, with parts and time recorded, is covered in our field service management guide.
The least discussed gain from this cycle is planning. When a failure is spotted while it develops, the work can be moved to a shift break or weekend when production stops anyway, spares arrive on normal lead times rather than by emergency order, and a prepared team does the job instead of a technician called out at night. So measure the value not only in avoided failures but also in the share of interventions that became planned. A mobile service report app that records fault types and readings consistently on every visit becomes one of the data sources for this analysis.
Where to start: a six-step pilot framework
- List your critical equipment. Rank machines whose failure stops the line, that have no standby and that take a long time to repair.
- Review two years of breakdown records. Which equipment failed, how often and why? If there are no records, that is a finding in itself.
- Choose one to three machines for the pilot. Equipment with a failure history and failure modes that show up in vibration or temperature is ideal.
- Define success up front. Measure unplanned downtime hours, emergency call-outs and similar metrics before the pilot starts, so there is a baseline to compare against.
- Bring the maintenance team into the process. Write down who checks what when an alert arrives and at what level a planned stop is scheduled; otherwise alerts get ignored.
- Scale on the evidence. Roll out to other lines for the equipment types where the pilot proved its value; for adding sensors to older machines, see retrofitting legacy machines.
Tracking downtime alongside OEE is the easiest way to translate gains into production terms; we show the calculation step by step in how to calculate OEE.
Common mistakes
When predictive maintenance projects fail to deliver, the cause is usually execution rather than technology. These are the mistakes we see most often on site:
- Putting sensors on everything. Covering non-critical equipment raises costs, multiplies alerts and scatters the team's attention. Keep the scope to machines whose failure is expensive.
- Setting alarms before learning normal behaviour. Every machine has its own vibration and temperature signature. Thresholds taken from a catalogue alarm either too late or too often; collect a few weeks of normal operating data first. Methods that learn normal behaviour and catch deviations are covered in our anomaly detection guide.
- Leaving alerts without an owner. If it is not written down who looks at an alert, how quickly, and when a planned stop is triggered, the system gets ignored after the first few false alarms.
- Separating data from maintenance records. If sensor data and the work actually done are not kept together, you cannot measure the model's accuracy or improve it over time. Record "what was found?" after every intervention.
- Scaling without measuring. If downtime hours and emergency call-outs were not measured before the pilot, nobody can say whether it worked. The roll-out decision should rest on comparison, not impressions.
- Leaving networking and security for later. Sensor gateways connect to the production network; which device reaches which network, and how, must be designed from the start (OT security for SCADA).
What these mistakes have in common is that the project should be managed as a change to the maintenance process, not as an equipment purchase.
The same predictive approach works for energy losses as well as machine failures. We cover catching leaks early from compressor and pipework data in our article on compressed air leak detection.
Usage-based maintenance works in hospitals as well as factories: servicing planned on real run hours means fewer medical devices pulled out of use just because the calendar says so. We explain the approach in our article on hospital medical equipment tracking.
How we handle this at Digital Bridge
We run predictive maintenance projects on a start-small, scale-on-evidence basis:
- A free site assessment. For Industry 4.0 projects we review your critical equipment list, breakdown history and existing infrastructure on site, at no charge.
- A pilot on one to three critical machines. We keep the pilot tight so that you see the system's value in concrete metrics within the first 30–60 days, and you decide on roll-out based on that data.
- One team from sensor to dashboard. Software and hardware come from the same team: where needed, our device manufacturing and IoT engineers handle electronics design, enclosures and embedded software, and technical feasibility for hardware is free.
- Analysis and alerts. We build the normal operating profile, track deviations with anomaly detection models and deliver alerts to your maintenance team by email, SMS or app notification.
- Integration with what you already run. We connect to our remote monitoring platform for dispersed sites, to MES on the production side, and to your maintenance management system.
For the equipment types we monitor and how the pilot runs, see our predictive maintenance page.
Cutting unplanned downtime enlarges usable capacity without new investment; our manufacturing capacity planning guide shows how that feeds into the calculation.
Next step
This week, sit down with your maintenance lead for half an hour and list the three breakdowns that stopped the line longest in the past year. For each one, ask: was there an early warning sign, such as noise, heat or rising current? If the answer is yes, that machine is your pilot candidate. Get in touch and we will define the pilot scope together during a site assessment.
To see how maintenance fits with production, quality and energy, browse our complete Industry 4.0 guide.