Skip to content
Reliability2026-08-11

Hidden Failure Modes: The Failures You Cannot See Until They Are Needed

Not every failure announces itself. A stuck-closed relief valve looks exactly like a working one until the vessel overpressures. Hidden failures affect protective and standby systems and need a fundamentally different strategy: a deliberate test, not condition monitoring. Why, and how to size the interval.

DA
Dzulfikar Ats Tsauri
Reliability Engineer
Share:

Not every failure announces itself. A pump that seizes makes noise, leaks, and stops pumping — everyone knows. A pressure relief valve that is stuck closed looks exactly like a pressure relief valve that works perfectly, right up until the moment it is needed and the vessel overpressures. A smoke detector with a dead battery shows no sign until there is smoke and no alarm. These are hidden failures — failures that produce no visible symptom under normal operation and are discovered only when a demand is placed on the failed component, or when someone tests it.

Hidden failures are the most dangerous class of failure in a plant, because the protective and standby systems they affect are exactly the ones that exist to prevent catastrophic events. And they require a fundamentally different maintenance strategy from the failures that show themselves. Run-to-failure and condition monitoring do not work on a hidden failure, by definition — there is nothing to monitor and no symptom to react to. The only defence is a deliberate test.

What Makes a Failure Hidden

A failure is hidden when there is no normal operating signal that reveals it. The component sits idle or in standby, doing nothing, until it is called upon. If it has failed silently in the meantime, the demand finds it dead. The classic examples:

  • Protective devices — pressure relief valves, burst discs, safety instrumented systems, emergency shutdown systems, overspeed trips, fire and gas detectors. Their job is to do nothing until a dangerous condition, then act instantly.
  • Standby equipment — the backup pump, the standby air compressor, the redundant cooling fan, the emergency generator. They run only when the primary fails.
  • Alarms and indicators — smoke detectors, level switches, vibration trips, low-lubrication-pressure trips. They signal only when a threshold is crossed.
  • Non-running redundancies — the second of two parallel pumps running in duty and standby, the spare that carries no load.

What unites them: under normal operation, both the healthy unit and the failed unit look identical. The failure is invisible until the unit is needed or tested.

Why the Usual Strategies Fail Here

The standard maintenance strategies are built for visible failures, and each one breaks down on hidden failures.

Run-to-failure fails immediately — you cannot "run" a standby pump to failure, because it is not running, and the failure only manifests when the primary fails and the standby is called upon, at which point you have lost both. (See run-to-failure for when RTF is and is not appropriate.)

Condition monitoring fails because there is usually nothing to monitor. A vibration sensor on a standby pump that is not turning reads zero whether the bearing is healthy or seized. There is no signal until the pump starts, by which point it is too late.

Time-based preventive maintenance — inspecting or overhauling on a calendar — works but is often wasteful, because it replaces or services units that are perfectly healthy on a fixed schedule regardless of actual condition.

The strategy that actually fits hidden failures is failure-finding: a deliberate, scheduled test that exercises the component and confirms it would perform its function if called upon.

Failure-Finding and Proof Testing

A failure-finding task (sometimes called a proof test, particularly in the safety-instrumented-system world) is a scheduled check that the hidden component actually works. The test is defined by the component:

  • A pressure relief valve is bench-tested or popped on a set schedule to confirm it lifts at the correct pressure.
  • A standby pump is test-run on a schedule to confirm it starts, reaches pressure, and runs without distress.
  • An emergency generator is load-tested weekly.
  • A safety instrumented system is fully proof-tested on the interval dictated by its safety integrity level, exercising the sensor, logic solver, and final element end-to-end.
  • A smoke detector is function-tested on the maintenance round.

The interval of the test is the critical decision. Test too often and you waste effort and add wear (each relief-valve pop shortens its life slightly; each standby-pump start adds a cycle). Test too rarely and the component sits failed for a long time before anyone knows, raising the average unavailability — the chance that it is failed at the moment it is needed. The right interval is calculated from the acceptable probability of failure on demand and the component's failure rate, and it is the answer to the question: how long are we willing to have this protection be dead without knowing?

The Math: Failure on Demand

The reason hidden failures need this treatment is the math of failure-on-demand. For a visible failure, the failure is detected immediately, so the downtime is just the repair time. For a hidden failure, the downtime stretches from the moment of failure to the moment of discovery — which, without a test, could be months or years (until the next demand). The average unavailability of an untested hidden component is roughly half its mean time between failures, which for a reliable component can still mean it is dead a substantial fraction of the time when the rare demand arrives.

Scheduled failure-finding caps this dead-time at the test interval. If you test monthly, the component can be dead for at most a month before you find out. If you test annually, it can be dead for up to a year. The test interval is therefore directly the maximum exposure window, and choosing it is choosing how much risk you accept. (This is the same logic that puts AND gates into a fault tree — see fault tree analysis — testing a protective layer is what makes it a real safeguard instead of an assumed one.)

Common Mistakes

  • Testing the easy half and not the whole function. Testing that a standby pump motor starts, but not that the pump reaches design pressure and flow. The start is not the function; the function is delivering the duty. Test the whole chain.
  • Testing under conditions that miss the failure. A relief valve tested at ambient temperature may lift correctly and still stick at operating temperature. The test must reproduce the demand conditions as closely as is safe.
  • No record of the test. A test that is not recorded did not happen, for audit and incident-investigation purposes. Every failure-finding task produces a dated record against the asset.
  • Ignoring the test results. If a test finds the component failed, that is a finding — root-cause it. A protective device that failed its proof test is telling you something about that device, that service, or that whole class of equipment, and ignoring the failure guarantees it will be discovered again, by the demand instead of by the test.
  • Treating redundancy as a substitute for testing. Two standby pumps are not twice as safe if neither is ever tested; they are two independent chances to be dead at the moment of demand. Redundancy reduces risk only when each unit is independently tested and maintained.

How OpexMX Supports It

OpexMX treats hidden-failure assets differently from run-to-failure and condition-monitored assets: each one carries a failure-finding task with a defined interval, a defined test procedure, and a recorded result. The interval is flagged when it is due and escalated when it is overdue, because an overdue proof test on a protective device is a silent gap in the plant's protection. Test results — pass, fail, and the as-found condition — are stored against the asset so that a pattern of proof-test failures surfaces as a bad-actor signal on that class of equipment. (See bad actor analysis.) And the test interval is visible alongside the asset's role in any fault tree, so the AND gates in a safety analysis are backed by real, recently-tested protective layers instead of assumptions.

Manage hidden-failure and protective-equipment testing in OpexMX →

Get maintenance insights in your inbox

Join operators getting practical CMMS tips, case studies, and product updates. No spam.