Skip to content
Reliability2026-08-11

Bad Actor Analysis: Which Assets Are Eating Your Maintenance Budget

A small minority of assets produces the large majority of failures, downtime, and cost. These are the bad actors. Find them, fix them properly, and a maintenance team can cut its reactive workload by a third in a single quarter. Here is how to identify and target them.

DA
Dzulfikar Ats Tsauri
Reliability Engineer
Share:

The Pareto principle is brutal in maintenance: a small minority of assets produces the large majority of the failures, the downtime, and the maintenance cost. These are the bad actors — the pumps that fail every quarter, the motors that trip weekly, the valves that stick open monthly. Find them, fix them properly, and a maintenance organisation can cut its reactive workload by a third in a single quarter without touching any other part of the program. Ignore them and the team drowns fixing the same machines forever while the rest of the plant gets no attention at all.

Bad actor analysis is the systematic identification of these few assets. It is the targeting step that decides where defect elimination, root cause analysis, and condition monitoring actually pay off — because none of those techniques are worth applying evenly across every asset. They are worth applying to the assets that fail the most.

What a Bad Actor Is

A bad actor is any asset that consumes a disproportionate share of reliability resources, measured three ways:

  • Failure frequency — number of failures or corrective work orders in a period. The asset that generates twelve corrective work orders a year on a plant where the median is one.
  • Downtime contribution — total downtime hours attributed to the asset. Some assets fail rarely but take the line down for a week each time; they are bad actors by impact, not count.
  • Maintenance cost — parts plus labour spent on the asset. The pump that eats three rebuilds a year in parts and twenty technician-hours a month in attention.

An asset can be a bad actor on any one of these axes. The worst are bad actors on all three. The point of the analysis is to rank every asset on each axis and look at the top of each list — the same names keep appearing.

The Pareto in Practice

Run the numbers on a typical plant and the curve is almost always the same shape. Sort assets by failure count descending and plot the cumulative percentage: the top 10 to 20% of assets typically account for 60 to 80% of the failures. The same shape shows up for downtime and for cost. This is not a coincidence or a sign of poor maintenance — it is the natural distribution of failure in complex systems, and it is the reason even-magnitude maintenance effort fails. Spending the same engineering attention on the bottom 80% of assets as on the top 20% is wasting four-fifths of it on assets that do not fail.

Bad actor analysis turns that Pareto curve into a target list. Take the top ten by failure frequency, the top ten by downtime, and the top ten by cost; merge the lists; and you have a working set of maybe twelve to fifteen assets that warrant concentrated reliability effort. Everything else gets the standard PM program and not a minute more.

Doing the Analysis

The analysis itself is a data exercise, and it lives or dies on work-order data quality.

  1. Pull corrective work orders over a meaningful period — twelve months minimum, twenty-four is better. Failures, not PMs. The signal is in the corrective log.
  2. Aggregate by asset. Count failures per asset. Sum downtime per asset. Sum parts plus labour cost per asset. Three rankings.
  3. Apply the Pareto cut. Take the top quintile (or the top 10 to 20 names) from each ranking. Merge.
  4. Sanity-check the list with the floor. The technicians know who the bad actors are before you run the numbers — if your list does not match their intuition, either the data is dirty or the analysis cut is wrong. Both are fixable; neither should be ignored.

The common data-quality failure is asset assignment. If work orders are raised against a generic "line 3" asset instead of the specific machine that failed, the analysis cannot see the bad actor — the signal is smeared across the parent. Clean asset-level assignment is a precondition. (This is the same data-quality discipline discussed in our CMMS data migration piece — bad actor analysis is the payoff for getting the asset register right.)

Why Bad Actors Repeat

Once the list is in hand, the next question is why each asset is on it. Bad actors are almost never bad luck. They fall into a few recurring root causes:

  • Wrong strategy. The asset is on run-to-failure when it should be on preventive, or on calendar PM when it should be on condition-based. Strategy mismatch is the most common cause. (See run-to-failure and the criticality matrix for the framework.)
  • Unfixed root cause. The asset fails, gets fixed at the symptom level, and fails again because nobody eliminated the cause. This is the failure mode that defect elimination exists to break.
  • Application mismatch. The asset is undersized, the wrong type, or operating outside its design envelope. A pump sized for a duty it no longer performs cavitates and eats seals. No amount of maintenance fixes a selection error.
  • Installation error. Misalignment, poor foundation, pipe strain, inadequate lubrication from day one. The asset was set up to fail and has been paying for it since commissioning. (See laser alignment.)
  • Operating abuse. Operators running the asset outside its design limits — over-speeding, dry-running, cycling it hard. The maintenance team fixes the damage; the operating practice recreates it.

Diagnosing which of these applies to each bad actor is where the real engineering happens. Run a root cause analysis on each top failure. (See root cause analysis and fault tree analysis.) The answer is rarely that the asset is just unreliable — it is a specific, fixable cause that the repeat-failure pattern was hiding.

The Fix and the Follow-up

Bad actor analysis without follow-up is a waste of a spreadsheet. Each asset on the list gets a targeted reliability project: a root cause investigation, a strategy correction, an application fix, or a condition-monitoring deployment. The project has an owner, a deadline, and a success metric — and the success metric is always the same: does the asset drop off the list next quarter?

Track the bad actor list over time. A healthy program watches the top names change quarter to quarter as old bad actors get fixed and a new set surfaces. A sick program watches the same names sit at the top year after year, because nobody followed through on the fix. The list itself is a leading indicator of whether the reliability program is actually working.

How OpexMX Supports It

OpexMX produces the bad actor list automatically — a Pareto of assets by corrective failure count, by downtime, and by maintenance cost, refreshed from the work-order history on demand. Each bad actor links straight to its failure history and work-order chain, ready for root cause analysis. Once a fix is applied, the asset's failure-rate trend before and after is plotted against it, so you can see whether it actually dropped off the list or just went quiet for a month. The result is bad actor analysis that turns into targeted reliability work and measures its own success, instead of a one-off slide that nobody acts on.

Pull your bad actor list in OpexMX and start fixing the right assets →

Get maintenance insights in your inbox

Join operators getting practical CMMS tips, case studies, and product updates. No spam.