Skip to content
Reliability2026-08-11

Run-to-Failure Maintenance: When Doing Nothing Is the Right Strategy

Run-to-failure is not neglect — it is a deliberate maintenance strategy for the right assets. What RTF actually requires, when it pays off, when it is reckless, and how to decide which assets earn it.

DA
Dzulfikar Ats Tsauri
Reliability Engineer
Share:

The phrase sounds like an admission of defeat: run the machine until it breaks. In a lot of maintenance cultures, "run-to-failure" is exactly what lazy or under-resourced teams get accused of. But run-to-failure (RTF) is not the absence of maintenance. Done deliberately, on the right assets, it is the cheapest correct strategy — and forcing preventive maintenance onto assets that do not need it is how reliability budgets get wasted.

The whole game is matching the strategy to the asset. RTF is one valid answer. It is just the wrong answer for a lot of the assets people quietly apply it to.

What Run-to-Failure Actually Means

RTF is a deliberate decision: we will not attempt to prevent this asset's failure. We will let it fail, then restore it. The decision is intentional, documented, and backed by preparation — a spare is on the shelf, the labour to replace it is available, and the consequences of failure are understood and acceptable.

That last clause is the entire gate. RTF is only correct when the cost and disruption of preventing failure exceed the cost of letting it happen. For a non-redundant pump on a critical line, that is never true. For a desk fan in the office, it almost always is.

What RTF Requires (That "No Maintenance" Does Not)

Calling something run-to-failure and then being surprised when it breaks is not a strategy — it is just poor planning. Genuine RTF has four preconditions:

  • A spare on the shelf or on short lead time. When it fails, you cannot afford a six-week parts wait. Safety stock is sized so the replacement is available now.
  • Known, fast repair path. Labour, tools, and procedure are ready. MTTR for that asset is predictable.
  • Failure consequence that is acceptable. The failure does not injure anyone, does not breach a safety or environmental limit, and the production or quality loss is bounded and cheap relative to the cost of preventing it.
  • Redundancy where the asset matters. If the function is critical but the individual unit is not, run two in parallel (one fails, the other carries load, you fix the failed one). This is the classic N+1 setup on cooling pumps, air compressors, and standby generators.

Skip any of these and you do not have RTF — you have an unplanned failure waiting to happen, with a fancy name attached.

Where RTF Pays Off

RTF tends to be the right call when one or more of these hold:

  • The asset is cheap and the failure is cheap. A twenty-dollar sensor, a small motor, a light fitting. The engineering hours to build a PM plan cost more than the asset.
  • There is functional redundancy. Two pumps, one running one standby, automatic switchover. Individual failure is invisible to production.
  • Failure is benign and obvious. The asset fails safe, the failure is immediately detectable, and there is no secondary damage. A desk lamp fails — you see it is off, you replace it.
  • Condition monitoring is uneconomic. The asset does not justify a vibration sensor, oil sampling, or even a weekly inspection round. The data costs more than the failure.
  • Failure is genuinely random and unpredictable. Some failure modes are not wear-out but random events (certain electronic failures, foreign-object damage). PM based on age does not help; readiness to respond does.

Where RTF Is Reckless

RTF is the wrong answer when:

  • The asset is a single point of failure for a critical function. No redundancy, no bypass. When it stops, the line stops, the plant stops, or safety is compromised.
  • Failure causes secondary damage. A bearing seizes and destroys the shaft, the coupling, and the motor windings. The ten-dollar bearing becomes a five-thousand-dollar rebuild.
  • Failure has safety or environmental consequences. A pressure relief device, a fire suppression component, an emergency stop. These fail silent and people get hurt. Never RTF.
  • Failure is hidden. The asset fails and nobody knows until a demand is placed on it — typically protective functions and standby equipment. These need failure-finding tasks, not run-to-failure.

The pattern: high consequence + no redundancy = never RTF. Those assets earn preventive or predictive maintenance.

How to Decide, Asset by Asset

This is what a criticality analysis is for. Score each asset on consequence (safety, environmental, production, cost) and on failure detectability, then route it:

  • High consequence, single point of failure → preventive or predictive maintenance.
  • Critical function but redundant → run-to-failure on the individual unit, with tested automatic switchover and a stocked spare.
  • Low consequence, cheap, benign failure → run-to-failure, with a min-max spare.

The point is that the decision is explicit and written down per asset, not applied by default to everything that did not get a PM schedule. (For the framework that produces these decisions systematically, see our piece on the equipment criticality matrix.)

The Hidden Cost of Over-Maintaining

The mistake that gets less attention than under-maintaining is over-maintaining. Tearing down a perfectly healthy pump every six months "just in case" costs labour, parts, and downtime — and every intervention is itself a chance to introduce an infant-mortality failure (reassembly error, wrong torque, contaminated lubricant). For an asset that would have run two more years untouched, the PM program made it worse, not better.

A mature maintenance program is not the one with the most PMs. It is the one with the right PMs on the right assets — and deliberate RTF on the rest.

The Inventory Angle

RTF shifts cost from labour (PM tasks) to inventory (spares). That is usually a good trade — a part on a shelf costs less than a preventive program — but it only works if the spare is actually there when needed. RTF without a stocked spare is not a strategy, it is a wish. Min-max levels on RTF spares, with automatic reorder, are the inventory counterpart that makes the strategy real. (See our optimal spare parts inventory guide for sizing those levels.)

How OpexMX Handles It

OpexMX lets you tag each asset with its maintenance strategy — run-to-failure, preventive, or predictive — and behaves accordingly. RTF assets are excluded from PM schedules but tracked for failure cost and repair MTTR, so you can see whether the strategy is still paying off. Spares for RTF assets sit in the parts register with their own reorder points, and a failure opens a work order against the right asset with the right spare pre-linked. Nothing is over-maintained, and nothing fails unprepared.

Map your asset strategies in OpexMX — book a walkthrough →

Get maintenance insights in your inbox

Join operators getting practical CMMS tips, case studies, and product updates. No spam.