Skip to content
Reliability2026-08-11

Defect Elimination: Getting Off the Fail-and-Fix Treadmill

Most maintenance teams are excellent at fixing things, and that is the wrong target. Defect elimination breaks the fail-and-fix loop by removing causes instead of speeding up repair. The process, the symptom-vs-cause skill, and why the no-time objection is backwards.

DA
Dzulfikar Ats Tsauri
Reliability Engineer
Share:

Most maintenance organisations are very, very good at fixing things. Give them a broken pump and they will have it running again by lunch, with a story about the time they did it faster under worse conditions. The problem is that being good at fixing is the wrong target. The right target is not having to fix the same thing again.

Defect elimination is the discipline that breaks the fail-and-fix loop. Instead of getting faster at repair, it asks why the failure happened and removes the cause, so that particular failure stops happening — to that asset, and often to every similar asset. It is the difference between a maintenance organisation that gets busier every year and one that gets quieter every year as the recurring failures disappear.

The Reactive Loop

The default state of most maintenance teams is a loop: an asset fails, the team responds, fixes it, closes the work order, and waits for the next failure. The work is urgent, visible, and heroic — and it is also a treadmill. Each fix resets the clock on the same failure mode that will recur, because the fix addressed the symptom, not the cause.

A pump's mechanical seal fails every four months. The team replaces it every four months. They are good at it — they have the part kitted, the procedure down, MTTR under control. And they will replace that same seal, on that same pump, four times this year, twelve times in three years, because nobody ever asked why the seal keeps failing. That is the reactive loop, and most plants are trapped in it without realising, because the heroics of the fix are more visible than the futility of the repeat.

Defect Elimination Breaks the Loop

Defect elimination inserts a different step between "fix" and "wait for next failure": understand the cause, and eliminate it. The seal that fails every four months — is it being installed wrong? Is the shaft deflected by misalignment or pipe strain? Is the flush arrangement wrong for the service? Is the seal itself the wrong type? Each is a fixable root cause. Find the right one, eliminate it, and that seal runs for three years instead of three months. One asset, eight avoided failures, one avoided rebuild. Multiply across the plant.

The discipline is not new technology. It is the systematic application of investigation methods that already exist — root cause analysis and fault tree analysis for the serious events, structured "five whys" for the routine ones — to recurring failures, and then following through to eliminate the cause rather than logging the fix. (See our primers on root cause analysis and fault tree analysis.)

The Process

A defect elimination program has four moving parts.

Capture defects at the point of detection. The defects that become recurring failures are often spotted early — by an operator on a round, a technician on a PM, a condition-monitoring alert — and then lost because there is no easy path from "I noticed something" to "it got analysed and fixed." The capture has to be frictionless: flag it on a mobile device at the moment it is seen, with a photo, against the asset. (This is the same defect-capture flow that makes 5S rounds effective.)

Analyse, do not assume. Each recurring defect gets a real root cause analysis, not a guess. The recurring seal failure is not "bad seals." It is a specific mechanical or process cause, and finding it requires investigation, not assumption. The common failure here is stopping at the first plausible answer — "the technicians are installing them wrong" — instead of drilling to the actual cause.

Eliminate the cause, then verify. Apply the fix to the root cause, and — the step most programs skip — measure whether the failure actually stopped. If the seal still fails every four months after the "fix," the root cause was wrong, and the loop restarts. Verification closes the loop honestly.

Feed the learning back. If the root cause was misalignment, every similar pump in the plant is a candidate for the same failure. Fix the family, not just the individual. This is where defect elimination compounds: each investigation prevents a class of failures, not just one.

Start With the Bad Actors

You cannot eliminate every defect at once, and trying to scatters attention. Start with the bad actors — the small set of assets that generate a disproportionate share of the failures and the maintenance hours. In most plants, a Pareto analysis shows that 20% of the assets produce 80% of the recurring failures. Those assets are where defect elimination pays back fastest, because each eliminated cause removes not one failure but a recurring stream of them. Pull the top ten bad actors from the work-order history, run a defect-elimination pass on each, and the reactive workload drops measurably within a quarter.

This is also why trustworthy failure data matters — you cannot find the bad actors if your work-order history is too dirty to aggregate. (The data quality discussion in our CMMS data migration piece applies here directly.)

Symptom vs Cause — the Core Skill

The skill that makes defect elimination work is distinguishing symptom from cause, and it is harder than it sounds because the symptom is always loud and the cause is always quiet.

  • Symptom: "the bearing failed." Cause: the lubrication schedule was wrong, or the alignment was out, or contamination ingress was uncontrolled.
  • Symptom: "the motor tripped on overload." Cause: the process was making it push past its rated load, or the cooling was blocked, or the winding was degrading.
  • Symptom: "the pipe leaked at the weld." Cause: vibration from an uncorrected resonance was fatiguing the weld.

The fix applied to the symptom — replace the bearing, reset the motor, reweld the pipe — guarantees recurrence. The fix applied to the cause eliminates it. The whole discipline is the refusal to stop the investigation at the symptom.

The Objection: "We Don't Have Time"

The standard objection to defect elimination is that the team is too busy fighting fires to step back and eliminate causes. This is exactly backwards. The team is too busy fighting fires precisely because it does not eliminate causes. Every defect eliminated is a fire that does not happen next month, and the month after, and the year after that. Defect elimination is the only path off the reactive treadmill; doing more reactive work faster is not.

The investment is front-loaded — the analysis takes time you do not have, in the first month — and the payback is compounding. A team that eliminates ten recurring causes this quarter fights fewer fires next quarter, which frees the time to eliminate the next ten. Within a year the reactive workload is materially down and the team is doing planned, eliminative work instead of emergency response. That is the mature state, and it is reachable, but only by refusing the argument that there is no time to start.

How It Connects to the Rest of Reliability

Defect elimination is not a standalone program — it is the execution arm of a reliability-centered approach. The analysis tools come from RCA and FTA. The defect capture comes from operator rounds and condition monitoring. The prioritisation comes from criticality analysis. The cultural engine is the same one that powers continuous improvement and TPM and reliability-centered maintenance. Defect elimination is what makes all of those real on the shop floor, by turning their frameworks into assets that stop failing.

How OpexMX Supports It

OpexMX surfaces the bad actors automatically — a Pareto of assets by recurring failure count and maintenance cost, refreshed from the work-order history — so a team knows where to point its elimination effort without a spreadsheet exercise. Each flagged defect captured in the field becomes a record linked to the asset, ready for RCA; each root cause, once found, can be attached to the asset's failure history so the pattern is visible across the family. Verification is built in: the failure-rate trend before and after the fix is plotted against the asset, so you can see whether the cause was actually eliminated or just hidden. The result is a program that measures its own success in failures that stopped happening.

Find your bad actors and start eliminating defects in OpexMX →

Get maintenance insights in your inbox

Join operators getting practical CMMS tips, case studies, and product updates. No spam.