7 Reasons Manufacturing Plants Underperform

Stagnation Slaughters. Strategy Saves. Speed Scales.

The standard advice when a plant misses is to run a root cause analysis on the worst line. Five whys, a fishbone, a cross-functional team, two weeks of meetings. I have sat through a lot of these and they almost always terminate in the same place: a machine, a material, or a person. Which is to say they terminate at the first visible thing, because that is what the method is built to find.

The trouble is that the visible things are downstream. A plant does not underperform because a machine is unreliable. It underperforms because it took nine days for anyone with authority to hear that the machine was unreliable, and by then the schedule had absorbed the loss and made it permanent. Root cause analysis run on a line finds line causes. The causes that matter live in the layer above the line, and nobody convenes a fishbone about those.

What follows is the taxonomy I use instead. Seven causes, split into the three that look like the problem and the four that actually are it. If you are working the full recovery sequence, this diagnostic sits inside stage one of how to turn around a manufacturing plant. Read this first, then go count.

Cause one: labor, which is almost never the cause

Labor is the most frequently named cause of plant underperformance and the least frequently correct one. It is named first because it is the most visible variable in the building and because it moves. Turnover is up, absenteeism is up, the new hires take longer to qualify, the experienced operators retired inside the same eighteen months. Every one of those observations is usually true. None of them explains the margin gap on its own.

What I look for is whether the labor symptom is stable or whether it tracks something else. In the plants where labor genuinely was the constraint, the pattern was flat: the plant could not staff a shift, full stop, and output was capped at the number of bodies present. That is a real and solvable problem, and it is solved with wages, scheduling, and qualification time rather than with anything in this article.

Far more often the labor symptom is oscillating, and oscillation means it is a response rather than a cause. Absenteeism climbs after a run of mandatory Saturdays. The mandatory Saturdays exist because the schedule slipped. The schedule slipped because a decision about a changeover sat unanswered for a week. Trace it backward two or three steps and you land somewhere in causes four through seven every time.

The field test is simple. Ask three operators on the worst line what would let them make more parts tomorrow. If they describe something about themselves, labor may be real. If all three describe something they have already told somebody about, you have a decision latency problem wearing a labor costume, and no amount of hiring will touch it.

Cause two: equipment, which is a symptom with a maintenance log

Equipment is the second most named cause, and it is unusual among the visible three because it comes with its own paper trail. Unplanned downtime is recorded. Work orders are recorded. Parts consumption is recorded. Which means that unlike labor, equipment can be checked rather than argued about, and the check usually reveals something more interesting than the machine.

Pull twelve months of unplanned downtime and sort it by asset. In a plant with a genuine equipment problem, the distribution is broad: many assets, all aging, all failing at rates consistent with their age and duty cycle. That plant needs capital and a maintenance strategy, and no operating change will substitute for either.

In most underperforming plants the distribution is not broad. It is two or three assets producing the majority of the downtime, and those assets have been on the capital request list for two or three cycles. At that point the cause is not the equipment. The cause is a capital allocation process that cannot distinguish between a machine that limits the constraint and a machine that does not, and that process lives in an office rather than on the floor.

Equipment is unusual among the visible causes because it comes with its own paper trail. The log will tell you whether you have a machine problem or a capital allocation problem.

The second check is the ratio of planned to unplanned maintenance hours. When unplanned dominates, the maintenance function has been converted into a reaction queue, and the reason is nearly always that planned maintenance windows keep getting surrendered to the schedule. That is a decision, made repeatedly, by someone. It is not an equipment fact.

Cause three: demand, which is sometimes real

Demand is the one visible cause that is genuinely the answer a meaningful share of the time, which is precisely why it needs the most rigorous test. Volume is down, the plant is running at 60% of what it was designed for, and every unit absorbs a larger share of a fixed cost base. That is a real and specific problem with a real and specific set of responses, none of which resemble the responses to an execution problem.

The test that separates it is whether the plant is losing orders or losing to orders. A plant losing orders is quoting and not winning, and the diagnosis moves to price, lead time, and position. A plant losing to orders is winning the work and failing to convert it: the backlog exists, the past-due list is long, and the constraint is inside the building.

Where this gets expensive is the response. Cost reduction applied to a demand problem shrinks the business permanently, because you remove the capability that would have served the recovery and you remove it faster than demand returns. I have watched competent operators do this with the best intentions, and the tell afterward is a plant that is profitable at a volume it will never grow past. The fork is important enough that I gave it its own piece: cost problem or demand problem.

One more note from the field. When a plant is losing orders, ask whether the losses cluster in specific product families rather than spreading evenly. Clustered losses are a position problem in one segment, which is narrow and fixable. Even losses across the whole book are a price or delivery problem, which is broad and slower.

Cause four: visibility latency

Visibility latency is the gap between when something goes wrong and when somebody who can act on it finds out, and it is the first of the four causes that actually produce underperformance. It is invisible in every report because reports are its symptom rather than its measure.

The mechanism is arithmetic rather than cultural. A shift loses an hour in the first two hours of a run. Nobody counts until the shift ends. The count is entered the following morning, aggregated into a weekly, and reviewed on Tuesday of the following week. By the time anyone competent looks at the number, nine days have passed, three more shifts have run the same way, and the hour is now nine hours that have been absorbed into the schedule and rebaselined as normal capacity.

What makes this cause so durable is that nothing about it looks like failure. Everyone did their job. The count was accurate. The report was on time. The review happened on schedule. The loss is structural, produced by the frequency of measurement rather than by anyone’s performance, and it will survive every improvement program you run on top of it.

The diagnostic is to compare measurement frequency against decision frequency for each role. A supervisor makes resource decisions hourly. If his number arrives daily, he is making twelve decisions blind for every one he makes informed. Plants that fix this see movement within two weeks from nothing but the change in frequency, which is the clearest evidence that the original problem was never on the line.

Cause five: accountability diffusion

Accountability diffusion is what happens when a number has three owners, which functionally means it has none. It is the quietest of the seven causes and the hardest to name in a meeting, because every individual involved is working hard on their piece and can prove it.

The pattern is recognizable. On-time delivery is owned by operations, planning, and customer service. When it misses, operations points at a schedule change, planning points at a capacity commitment, and customer service points at a promise date it was given rather than one it made. All three accounts are accurate. The number does not move, and it will not, because improving it requires someone with the authority to trade one function’s convenience for another’s, and no such person is in the room.

What I look for is whether a number has a single name attached and whether that name has the authority to change the inputs. Ownership without authority is not accountability, it is exposure, and people who are exposed rather than accountable spend their energy building the record that will protect them rather than moving the number.

Ownership without authority is not accountability, it is exposure. People who are exposed rather than accountable spend their energy building a record instead of moving the number.

The related failure is the metric that is owned at the wrong altitude. Plant-level on-time delivery owned by the plant manager is not actionable, because it aggregates across product families with entirely different failure modes. Split it by family, attach each split to the person who controls that family’s schedule, and the diffuse problem resolves into three specific ones with three specific owners.

Cause six: decision latency

Decision latency is the median time between a question leaving the floor and an answer coming back, and it is the single most reliable predictor of plant performance I have found. It is also the least measured number in manufacturing, which is a strange fact about an industry that measures nearly everything else to four decimal places.

The mechanism compounds in a way the other causes do not. A supervisor waiting on an answer does not stop working, he works around, and the workaround has a cost: an extra changeover, a partial run, a batch made to the wrong sequence, a part expedited. Every workaround is individually rational and collectively expensive, and because none of them appears as a variance anywhere, the total cost of a slow decision layer is never assembled into a number anyone sees.

The second-order effect is worse. Supervisors calibrate on the observed return time. When escalation reliably takes nine days, they stop escalating and start absorbing, and the plant loses its early warning system entirely. At that point the general manager’s reports get quieter and more optimistic while the underlying performance degrades, which is a specific and dangerous condition because it looks like stability.

Measuring it takes a week. Log every decision that leaves the floor: what was asked, who it went to, when it returned. You will find a median that surprises you, and you will find that the destination is concentrated in one or two people who accumulated approval rights over a decade and who are now, without anybody intending it, the plant’s actual constraint.

Cause seven: portfolio drag

Portfolio drag is the accumulated cost of making things the plant should not be making, and it is the cause with the largest financial weight and the lowest diagnostic visibility. It hides because complexity cost sits inside overhead, and overhead gets allocated evenly across everything, so the standard cost system reports each product as roughly as profitable as its neighbors.

The costs that do not get attached are the ones that vary by product rather than by volume: changeover time, tooling upkeep, the inventory a low-runner forces you to carry, the forecast error it generates, the documentation it demands, and the share of engineering and customer service attention it absorbs. Attach even a rough estimate of those and the profitability distribution changes shape entirely. It stops being a slope and becomes a cliff with a long negative tail.

Complexity cost sits inside overhead, and overhead gets allocated evenly, so the standard cost system reports every product as roughly as profitable as its neighbor.

The operational symptom is a plant that feels busy and posts poor margins. Changeovers dominate the schedule, the constraint runs at low utilization because it is switching rather than making, planning spends its week resequencing, and everyone in the building can tell you the plant is working hard. It is. It is working hard on the wrong mix, and no amount of execution improvement will fix a mix problem.

The quick diagnostic before any full costing exercise: count the product families that account for the bottom decile of volume, then count the share of total changeovers they consume. When a small share of volume is consuming a large share of the constraint’s available time, portfolio drag is your cause, and it is the one whose repair pays the most.

Monday morning, log every decision that leaves your worst-performing line for one week. Nothing else. That single log will tell you which of these seven you are actually dealing with.

About the Stagnation Assassin

Todd Hagopian is a Fortune 500 transformation executive who has generated $3B+ in shareholder value across Berkshire Hathaway, Illinois Tool Works, Whirlpool, and JBT Marel, where he serves as VP of Global Product Strategy. Known as The Stagnation Assassin, he is the author of two published books: The Unfair Advantage: Weaponizing the Hypomanic Toolbox and Stagnation Assassin: The Anti-Consultant Manifesto. His blog is published in 15+ languages and read by operators worldwide. Bring him to your stage via his speaking page or connect with him on LinkedIn.

One question to answer yourself first: how many days does it take for a problem on your worst line to reach somebody with the authority to fix it? If you cannot answer that in a number, that is your cause. When you have the number and want to know what to do with it, tell me what it is.