The flawed prioritization: Failure frequency does not correlate with failure cost; downtime duration and response readiness drive the real damage.
- The outage trigger: A specialized two-cent bolt, ordered from a specific supplier, kept a machine down for weeks — the costliest failure that year despite being the cheapest and rarest.
- The structural blind spot: Predictive systems are built around expected failures, making them blind by design to rare ones never present in training data.
- The recommended audit: Manufacturers should identify uninstrumented parts dismissed as insignificant and flag any that couldn't be quickly replaced if they failed at the worst moment.
- The balanced takeaway: Monitoring common failures is still necessary, but covering both common and rare failures is essential to preventing significant downtime.
A single bolt costing about two cents took a factory down for the better part of a month. Not because it failed often — it rarely did — but because no one had ever thought it worth watching.
"It wasn't one of the failure modes that anyone had thought to instrument," writes Michael Podgortsev, Director of Data and AI Strategy and former CTO, in a Sept. 22, 2026 IndustryWeek column.
A Failure Nobody Modeled
Podgortsev describes a scenario in which a predictive maintenance system performed exactly as designed — monitoring wear parts, learning the patterns that precede common failures, catching them before they cost production time. Then a small bolt broke. There was no sensor on it, no data stream, no entry in the model. The system couldn't have seen it coming; the failure lived entirely outside what it was built to watch.
The damage compounded from there. The bolt was a specialized part, not something sitting in a drawer. It had to be ordered from a specific supplier, and the machine sat dead while everyone waited for delivery. Weeks passed before the part arrived and a technician installed it.
"A two-cent component took a production line offline for weeks and cost a fortune in lost output," Podgortsev writes. "The cheapest, rarest failure in the whole system turned out to be the most expensive one that year."
Frequency Is the Wrong Filter
Podgortsev's central argument: the instinct to prioritize monitoring based on how often something fails is backwards. Rare failures never appear in the training data, so a system tuned for the average, expected breakdown is blind to them by design. They don't show up as warnings — they show up as a dead machine and an alert that never fired.
"The cost of a breakdown has almost nothing to do with how often it happens," he notes. The real damage lies in how long the line stays down and how ready the response is.
His prescription is a different audit question. Not "Are we catching the failures we know about?" — most plants probably are — but "What could take a line down that we have no sensor for, and no plan to respond to?" He advises walking the machine and identifying uninstrumented parts that seemed too small or too rare to bother with, then asking which of them couldn't be fixed quickly at the worst moment. That intersection — invisible failure plus slow recovery — is where the next extended outage hides.
"The machine that's most likely to blindside you isn't the one failing in ways you understand," Podgortsev concludes. "It's the one about to break in a way your system was never built to see."
None of this invalidates conventional predictive maintenance. Monitoring common failures remains necessary and earns its keep. But a system optimized only for the expected failure leaves the door open to the one nobody thought to watch — and on a factory floor, that exposure is measured in weeks of downtime.
Is this your company?
This article features your business. Claim it to add your logo, contact details, and a link to your website — or upgrade to reach more buyers.
Did you know 80% of Press Releases trigger AI content warnings? Reach out and the M4S team can assist.
