Skip to content
Reliability2026-09-18

From Symptom to Root Cause: How AI Traced a Molding Machine's Rejects to a Compressor Bearing

Maintenance fixed the molding machine three times. The rejects came back within days every time. Here is how an AI evidence funnel walked from a one-line operator complaint to a worn compressor bearing two assets away — and put a price on fixing it before it failed.

OT
OpexMX Team
Share:

The machine was not broken. That was the uncomfortable part.

Over two weeks, the reject rate on molding machine M-14 at a plastics packaging plant climbed from 0.8% to 4.6% — mostly short shots and out-of-spec dimensions, worst during the mid-afternoon demand peak. Maintenance responded the way good maintenance teams do: three interventions in twelve days. A seal kit on the injection unit. A recalibration of injection pressure and hold profile. A clamp force tune-up.

Each fix worked. For a few days. Then the rejects crept back.

Everyone had a theory — mold wear, resin batch, ambient humidity. The supervisor finally typed the question into OpexMX the way the operator had been saying it on the floor all week: "mesin ini reject lagi, kenapa ya?" — this machine is rejecting again, why? What came back was not a search result. It was a structured investigation that ended at a screw compressor nobody had complained about, two assets away on the air side of the plant.

Why Three Good Fixes Could Not See the Cause

Single-asset troubleshooting has a blind spot, and M-14 found it. The reject signature was intermittent and load-correlated — it peaked exactly when the plant did. Every intervention targeted the machine itself, because the machine is where the symptom lived.

The actual pattern needed three data silos joined: reject timestamps from quality logs, the plant-air header pressure trace, and the compressor's vibration spectrum. No human had a reason to join those at two in the afternoon, let alone at two in the morning. Compressors are "someone else's asset" until they are everyone's problem. Cross-asset causality is precisely where tribal knowledge fails — and where a system that can see all three streams at once stops guessing.

Start From the Complaint, Not From a Form

RCA quality is capped by the quality of the problem statement, and operators speak symptoms, not failure modes. Forcing "kenapa ya?" into a dropdown menu loses the signal before the analysis starts.

So the copilot's first job is translation, not interrogation. From the complaint and the context it can already see — the asset it was typed against, recent work orders, the quality trend — it drafts a structured problem statement: asset M-14, symptom rising reject rate (short shots plus dimensional drift), timeframe two weeks and worsening, correlated condition mid-afternoon demand peak. The supervisor corrects one detail and confirms the rest. Then the funnel starts.

The Evidence Funnel: Four Passes, One Thread

The RCA is not a magic answer machine. It is a funnel, and every pass is auditable — each narrows the candidate list using evidence the plant already has on file.

Pass 1 — Fishbone (6M). Candidates are generated across Man, Machine, Method, Material, Measurement, Milieu: 23 in total, drawn from the asset's history, OEM documentation, and current process parameters. Nothing is discarded yet; the point of the fishbone is breadth before judgment.

Pass 2 — FMEA ranking. Each candidate gets scored by RPN — severity × occurrence × detection — using the plant's own failure mode library plus actual occurrence data from work order history. Twenty-three candidates collapse to seven above the action threshold. Resin humidity ranks mid-pack. Mold wear and air supply rank at the top.

Pass 3 — Fault tree, confirmed or refuted by evidence. For each surviving candidate, a small fault tree is built and every branch is checked against data rather than opinion. The mold wear branch: cavity pressure curves normal, mold inspection six weeks ago clean — refuted as the primary cause. The injection pressure branch: setpoints unchanged for months, controller logs stable — refuted. The air supply branch: the header pressure trace shows a 0.4 bar sag during exactly the afternoon window when rejects spike — confirmed. (A cross-encoder reranker searching past failures in the knowledge base had already surfaced two similar cases from other lines — the same complaint pattern, the same air-side resolution — which is what directed attention to the header trace this early.)

Pass 4 — 5-Whys. Rejects → short shots and dimensional drift → ejection and clamp pneumatics running unstable → header pressure sag at peak demand → compressor C-02 unloader cycling abnormally short → worn drive-end bearing changing rotor dynamics. Every "why" cites its evidence, and the whole chain is reviewable by a human before anyone touches a wrench.

Each pass is grounded by retrieval from the OEM manuals — compressor unloader specifications, the molding machine's air quality requirements — reranked by relevance, so the conclusions cite document sections, not plausible prose.

The Root Cause Was Two Assets Away

C-02 is the screw compressor serving the molding bay, and three weeks of its vibration data told the whole story. The BPFO envelope peak at the drive-end bearing had tripled. Overall velocity was at 7.8 mm/s RMS — past the 7.1 mm/s line into ISO 10816 zone D for this machine class, with BPFI, BSF, and FTF markers confirming the outer-race mode. (The mechanics of reading these peaks are in our vibration monitoring guide.)

Regime detection had flagged the same week before anyone connected it: a PCA model over the compressor's multivariate data marked discharge temperature up 6°C and unloader cycles 40% shorter than baseline — the machine working harder to hold the same header pressure. A degraded bearing changes rotor dynamics; volumetric efficiency drops; header pressure sags under peak demand; and the molding machine's pneumatics starve at exactly the moment the plant is running hardest.

The Weibull fit on the degradation trend turned the diagnosis into a deadline: 11 days of remaining useful life, 95% confidence interval 6 to 23 days. Not "a bearing is wearing" — a countdown. (How failure history becomes those confidence bands is covered in our Weibull analysis article.)

The Economics: $2,480 Planned vs $30,561 Failed

Detection changes maintenance from calendar-driven to condition-driven. The money question — what should we do about it, and when — is where prescriptive ranking earns its keep. For every candidate action, the engine computes a total expected cost: action cost plus residual failure risk.

ActionCost nowFailure riskTotal expected cost
Swap bearing kit in Wednesday's changeover window$2,480negligible (OEM procedure, new part)$2,480
Defer to the next planned line stop (19 days out)$0~87% chance of running to failure first (Weibull fit)≈$26,600 expected; $30,561 when it happens

The failure side itemizes like this: 7.2 hours of unplanned stop across molding and packaging at $3,400/hour ($24,480), accumulated scrap and re-inspection ($2,150), expedited bearing plus emergency courier ($1,900), overtime recovery crew ($1,281), and post-repair QA revalidation ($750). That is $30,561 — against a $2,480 planned swap, a twelve-to-one decision before probability-weighting even enters.

Even the $2,480 was a decision, not a quote. The engine checks stock first: no kit on the shelf. Then it ranks vendors by one rule — cheapest that delivers before the window. The cheapest kit quotes a 14-day lead time against an 11-day RUL median, so it is excluded no matter how attractive the price. The next vendor delivers in 3 days for $120 more — inside the window. The expedite courier is held as a fallback if the primary slips.

The maintenance manager approved the swap. It took four hours inside an already-scheduled changeover. The reject rate returned to baseline within one shift, and the compressor went back to ISO 10816 zone A/B — where it should have stayed all along.

Closing the Loop

The story does not end at the fix, because a fixed machine that teaches the organization nothing is only half a repair. The full loop runs: detection → work order → checklist execution → RCA → knowledge capture. The RUL countdown auto-opened the work order with the deadline attached; the checklist ran on mobile during the swap; the complete investigation — fishbone, fault tree, 5-Whys chain, cost comparison — was written back against the work order; and the knowledge base indexed the pattern: molding rejects + header pressure sag + compressor BPFO growth.

The next plant that hits this pattern starts from a confirmed failure story instead of a blank fishbone. (The strategy behind that loop is in our closing the loop article, and the fundamentals in what is root cause analysis.)

Worth stating plainly what the AI did not do: it did not replace the investigation — it ran it. Every pass was reviewable, every conclusion cited evidence a technician could check with a handheld. The human made exactly one decision, and it was the one humans should make: swap or defer, priced.

How OpexMX Handles RCA

OpexMX connects the whole loop end to end: vibration-based fault detection (BPFO/BPFI/BSF/FTF peaks, ISO 10816 zoning), Weibull RUL with confidence bands, an AI copilot that runs the evidence funnel from a complaint as informal as "mesin ini reject lagi, kenapa ya?", and prescriptive recommendations that rank actions by total expected cost — with vendor, lead time, and stock logic attached to every parts line.

The symptom can stay messy. The investigation won't be.

See how OpexMX turns symptoms into root causes — and root causes into priced decisions →

Get maintenance insights in your inbox

Join operators getting practical CMMS tips, case studies, and product updates. No spam.