Predictive Maintenance ROI: Where to Spend

Stagnation Slaughters. Strategy Saves. Speed Scales.

Executive summary: Predictive maintenance pays for itself only where downtime costs you output. On a constraint, an hour of unplanned downtime is throughput the system never recovers, so preventing it converts almost directly into money. On a machine with spare capacity, preventing the identical hour returns nothing to the system. This guide shows how to calculate the return honestly, why most maintenance budgets are pointed at the wrong assets, and how to sequence a program that pays back in weeks.

What is predictive maintenance?

Predictive maintenance uses condition data such as vibration, temperature, oil chemistry, and motor current to detect developing failures before they stop a machine, so intervention happens on a planned schedule. It sits between reactive maintenance, which waits for failure, and preventive maintenance, which replaces parts on a calendar regardless of condition.

The three approaches are worth separating clearly, because most plants run a confused mixture of all three and call it a strategy.

Reactive means run to failure. The machine stops, you fix it, production waits. It is the cheapest approach right up until the moment it is catastrophically expensive, and its cost is entirely determined by which machine failed.

Preventive means scheduled replacement on a time or cycle basis. Change the bearing every six months whether it needs it or not. This prevents many failures and wastes a great deal of remaining component life, and it still misses the failures that do not follow the calendar.

Predictive means watching actual condition and acting on evidence of degradation. You replace the bearing when the vibration signature says it is deteriorating, not when the calendar says so, and you schedule that replacement into a window where it costs you nothing.

Industry reporting on predictive programs commonly describes unplanned downtime reductions on critical assets in the neighborhood of a quarter. That figure gets quoted constantly in vendor material, and it is close to useless on its own, because it says nothing about whether the downtime you prevented was costing you anything. That is the question this whole discipline turns on, and almost nobody asks it before signing the purchase order.

Why does placement decide the return?

Because downtime only costs output where output is constrained. An hour lost on the constraint is an hour of system throughput gone permanently. An hour lost on a machine with spare capacity is absorbed by that spare capacity and costs the system nothing. Identical technology, identical downtime reduction, radically different return.

I have led transformations at Berkshire Hathaway, Illinois Tool Works, and Whirlpool, and this is the single most common misallocation I encounter in maintenance spending. Plants deploy sophisticated vibration monitoring and thermal imaging across their newest, most expensive equipment, because that is where the capital is and that is what feels worth protecting. Meanwhile the thirty-year-old machine that actually determines system output runs to failure, because it is old anyway.

That is exactly backwards. Your newest machine can break down with no consequence to the business at all if it is not the constraint, because the spare capacity absorbs it. Your constraint cannot break down without immediate, unrecoverable damage to what the company earns that day. Age, cost, and sophistication are irrelevant to this calculation. The only question that matters is whether the asset sets your output.

Run the comparison honestly and the asymmetry is uncomfortable. Take a program that eliminates 280 hours of unplanned downtime per year. Deployed on the constraint at a plant where a constraint hour generates roughly ten thousand dollars of throughput, that is 2.8 million dollars of recovered output annually. Deployed on a non-constraint with spare capacity, the same 280 hours recovered produces a nicer availability number on a report and precisely zero additional system throughput, because the system was never limited by that machine.

Identical predictive maintenance program, three different assets, three different returnsSame program. Same hours saved. Three returns.280 hours of unplanned downtime eliminated on each assetNewest assetnot the constraint$0 system throughputMost expensive assetnot the constraint$0 system throughputThe constraint30 years old, sets output$2.8MBudget follows capital value. It should follow constraint status.

How do you calculate predictive maintenance ROI?

Measure current unplanned downtime hours on the constraint, estimate the percentage the program will eliminate, multiply by throughput value per constraint hour, and compare against program cost. On a genuine constraint the payback is usually measured in weeks, which is why the calculation is worth doing before the vendor does it for you.

Work it through on a real structure.

Step one, establish constraint running hours. A three-shift operation running roughly 7,000 production hours per year gives you the denominator for everything else.

Step two, measure actual unplanned downtime. Not the reported figure, the measured one. Say unplanned downtime runs at six percent of running time. That is 420 hours a year the constraint is stopped for reasons nobody scheduled.

Step three, estimate the reduction. A well-run predictive program focused on a single critical asset can take a substantial share of that out. In one program I ran, where maintenance reviewed sensor data in short daily meetings and scheduled interventions into breaks, unplanned constraint downtime fell 67 percent over six months. Applied to 420 hours, that recovers roughly 281 hours.

Step four, value the hours. At approximately ten thousand dollars of throughput per constraint hour, a common figure once you compute price minus truly variable cost, 281 recovered hours is about 2.8 million dollars of annual throughput.

Step five, compare to cost. Sensors, analytics, and the labor to run the program on one critical asset typically land well under two hundred thousand dollars in year one and considerably less thereafter. Against 2.8 million dollars of recovered throughput, payback arrives inside the first month.

A constraint running 7,000 hours a year with 6 percent unplanned downtime loses 420 hours annually. Cutting that by 67 percent recovers roughly 281 hours. At about $10,000 of throughput per constraint hour, that is close to $2.8M a year, against a program cost that rarely exceeds $200,000 in year one.

Two cautions on this arithmetic. First, throughput value per constraint hour must be calculated as price minus truly variable cost, not as revenue and not as accounting margin. Using a loaded margin figure will understate the return badly, because most of what gets loaded into it does not change when the constraint runs an extra hour.

Second, the reduction percentage should be your own measured result after a pilot, not a vendor benchmark. Run it on one asset, measure honestly for a quarter, and use your own number to justify expansion. Nobody has ever regretted having their own data in that conversation.

Which monitoring technologies are worth deploying?

Four cover most industrial failure modes: vibration analysis for rotating equipment, thermal imaging for electrical and friction faults, oil analysis for gearboxes and hydraulics, and motor current signature analysis for drives. Start with whichever matches your constraint’s dominant failure mode rather than buying a full suite at once.

Vibration analysis

The workhorse for anything rotating: motors, pumps, fans, gearboxes, spindles. Bearing degradation, imbalance, misalignment, and looseness all produce characteristic signatures well before failure. If your constraint contains significant rotating equipment, this is almost always the first sensor to deploy.

Thermal imaging

Catches electrical connection faults, overloaded circuits, failing bearings, and friction problems. It is fast, requires no contact, and a single handheld camera on a route covers a lot of ground cheaply. It also doubles as an energy inefficiency finder, which occasionally pays for the camera by itself.

Oil analysis

Wear particles, contamination, and lubricant degradation show up in oil chemistry long before they show up as noise or heat. Essential for gearboxes, hydraulics, and large drives. The lab turnaround makes it a periodic rather than continuous method, which suits its role.

Motor current signature analysis

Reads electrical signatures to detect rotor bar damage, eccentricity, and load anomalies without touching the machine. Useful where physical sensor placement is awkward or where the failure mode is electrical rather than mechanical.

The temptation is to buy everything and instrument broadly. Resist it. Identify the two or three failure modes that have actually stopped your constraint over the past two years, deploy the technology that detects those specific modes, and expand only when the first deployment has proven itself. Sensors are cheap now. Analytical attention is not, and a program generating alerts nobody has time to review is worse than no program, because it trains people to ignore alarms.

How do you sequence a program?

Sequence it in four moves: confirm which asset is the constraint, review its actual failure history, instrument the specific failure modes that history reveals, then build the daily review and scheduling routine that turns alerts into planned interventions. The routine matters more than the technology and gets built last by most teams.

Move 1: confirm the constraint

Everything downstream depends on this. Instrumenting the wrong asset produces a technically excellent program with no financial return. Confirm through inventory accumulation, cycle time against takt, and availability data before you spend anything on sensors.

Move 2: review failure history

Pull two years of work orders and downtime records for that asset. What actually failed, how often, and how long did each event stop production? This tells you which failure modes to instrument and gives you the baseline you will measure improvement against.

Move 3: instrument the specific modes

Deploy the monitoring technology matched to the failure modes the history revealed. Set alert thresholds conservatively at first and tune them as you learn what normal looks like on that specific machine. Expect a period of false alarms while thresholds settle.

Move 4: build the review and scheduling routine

This is where programs live or die. Maintenance reviews sensor data in a short daily meeting, decides what needs intervention, and schedules that intervention into a window where the constraint is not producing. Without this routine you have bought a data collection system, not a maintenance program.

When should the intervention actually happen?

Never during scheduled constraint production. Predictive maintenance earns its return by converting unplanned stoppages into planned ones, and a planned stoppage that still happens during production hours has captured only part of the value. Schedule into breaks, shift changes, and planned non-production windows.

This point sounds procedural and it is worth a great deal of money. The whole value proposition of predictive maintenance is advance warning. Advance warning is only valuable if you use it to move the work to a moment when the machine was not going to be producing anyway.

I have run programs where maintenance received an automatic alert at the first signs of degradation and then intervened during lunch or shift change rather than during production. That discipline is what turned condition data into recovered throughput. The alternative, which I have also watched, is a plant that detects the developing failure accurately, congratulates itself on the detection, and then stops the constraint for two hours in the middle of a shift to fix it. The detection was excellent. The value capture was partial at best.

The rule I use for constraints is uncompromising. The constraint should never be idle during scheduled production. If it needs adjustment or a changeover, schedule it during a break or shift change. If it needs material, that material arrives five minutes early rather than five minutes late. If it needs maintenance, it happens in a planned window. Predictive data is what makes that level of precision operationally possible rather than merely desirable.

Predictive maintenance converts unplanned stoppages into planned ones. If the planned stoppage still lands in the middle of a production shift, you captured the detection and threw away most of the value. Schedule every constraint intervention into breaks, shift changes, or planned non-production windows.

What are the most common predictive maintenance mistakes?

Four dominate: instrumenting assets by capital value instead of constraint status, buying a full technology suite before proving one, generating alerts nobody has the routine to review, and intervening during production hours so the advance warning is wasted.

Mistake 1: budgeting by asset value

The default is to protect the newest and most expensive equipment. That instinct is financially sensible for asset preservation and irrelevant for throughput. Put the predictive maintenance budget where your constraint is, not where your newest machines are. A failure on a non-constraint may cost you a repair. A failure on the constraint costs you the day’s output.

Mistake 2: buying the full suite first

Vendors sell platforms. Plants buy vibration, thermal, oil, and current monitoring simultaneously across dozens of assets, then discover they lack the analytical capacity to act on any of it. Start with one asset and the failure modes its own history shows. Prove the return, then expand with your own numbers behind you.

Mistake 3: alerts without a routine

The technology generates signals. If no one owns a daily review and no mechanism converts a signal into a scheduled intervention, the signals accumulate and get ignored. Within months the alerts are noise and the program is shelfware. Build the routine before you expand the sensors.

Mistake 4: measuring the wrong outcome

Programs get reported on availability improvement or alert counts, neither of which is money. Report the outcome in constraint hours recovered and throughput value. That framing also protects the program budget, because it puts the return in terms a CFO recognizes rather than in terms only maintenance cares about.

My own mistake, made more than once, was underestimating how long threshold tuning takes. The first six to eight weeks of a new deployment produce false alarms while you learn what normal looks like on that specific machine. If you have promised results in month one, those false alarms will be used as evidence the program does not work, and you will lose the political room to finish tuning it. Set expectations for a tuning period up front. It is a small piece of change management that protects an otherwise sound program.

Predictive maintenance: operator FAQ

How do you calculate predictive maintenance ROI?

Measure actual unplanned downtime hours on the constraint, estimate the share the program will eliminate, multiply by throughput value per constraint hour, and compare to program cost. For example, cutting 420 annual downtime hours by 67 percent recovers about 281 hours, worth roughly $2.8M at $10,000 per constraint hour.

Where should a predictive maintenance program start?

On the constraint, regardless of that asset’s age or cost. Downtime only reduces output where output is limited, so preventing an hour of downtime on a machine with spare capacity returns nothing to the system. Budget by constraint status rather than by capital value, which is how most plants allocate it.

What is the difference between preventive and predictive maintenance?

Preventive maintenance replaces components on a time or cycle schedule regardless of their actual condition, which wastes remaining life and still misses failures that do not follow the calendar. Predictive maintenance monitors real condition through vibration, temperature, oil chemistry, or motor current and intervenes on evidence of degradation.

Which monitoring technology should you deploy first?

Whichever matches the failure modes that have actually stopped your constraint over the past two years. Vibration analysis suits rotating equipment, thermal imaging catches electrical and friction faults, oil analysis covers gearboxes and hydraulics, and motor current analysis reads electrical signatures. Start with one, prove it, then expand.

About the Stagnation Assassin

Todd Hagopian is a Fortune 500 transformation executive who has generated $3B+ in shareholder value across Berkshire Hathaway, Illinois Tool Works, Whirlpool, and JBT Marel, where he serves as VP of Global Product Strategy. Known as The Stagnation Assassin, he is the author of two published books: The Unfair Advantage: Weaponizing the Hypomanic Toolbox and Stagnation Assassin: The Anti-Consultant Manifesto. His blog is published in 15+ languages and read by operators worldwide. Bring him to your stage via the speaking page or connect with him on LinkedIn.

Next step: a maintenance placement review

Your best monitoring is probably on your newest machine, and your output is probably set by an old one nobody watches. Book a 20 minute review and I will help you find which asset actually sets your throughput, then show you what the downtime on it is costing you every year. Start the review here.