Monitor vs Replace on an Intermittent: Decision Tree
Why this matters
An intermittent fault forces a bet. You either replace a part you cannot prove is bad, or you walk away and tell the customer to call when it acts up again. Bet wrong on replace and you charge for a part that does not fix it, then eat the callback. Bet wrong on monitor and the customer sits in the cold waiting for a failure you could have caught. The good techs do not flip a coin here. They have a rule for which way to lean, and they say it out loud to the customer so the decision is shared. This tree is that rule.
Start here: rule out a hazard before you choose
If the intermittent has any safety dimension, the monitor-vs-replace question does not apply yet. Make it safe first - kill power, shut the gas or water, verify dead or depressurized. A breaker that trips at random, an occasional burning smell, a heat source that "sometimes" overheats, or anything mixing water and electricity is not a part to monitor in service. Stabilize the hazard, then decide whether the component gets replaced now or the unit stays locked out until it is proven safe. Everything below assumes a fault with no safety consequence.
The two failure modes you are choosing between
- Replace too eagerly wastes the customer's money on good parts and erodes trust when the fault returns. It is the lazy answer that feels like progress.
- Monitor too long leaves the customer exposed to a failure that matters, and turns a fixable problem into a string of unbillable return trips.
Neither is free. The whole art is matching the choice to the evidence and the consequence.
The four inputs to the decision
Score the call on four axes before you lean either way:
- Strength of evidence. Did one component clearly move the reading when you stressed it, or is everything passing? Strong evidence pushes toward replace; clean tests push toward monitor.
- Consequence of recurrence. If it fails again, is it an inconvenience or a real problem (no heat in winter, no water, a flooded floor, a freezer full of food)? High consequence pushes toward replace.
- Cost and reversibility of the part. A low-cost, easily swapped part you can do today is cheap to gamble on. A major, labor-heavy component is not a guess you make on suspicion.
- Pattern. Is there a predictable trigger you can instrument and wait out, or is it truly random? A known pattern favors monitor-with-a-recorder; pure randomness with high consequence favors a targeted replace.
The decision table
| Evidence | Consequence | Part cost | Lean |
|---|---|---|---|
| Strong (one part moved the reading) | Any | Any | Replace that part, then verify |
| Weak / clean tests | High (no heat, no water, flooding) | Low, quick | Replace the single most-likely item, instrument, watch |
| Weak / clean tests | High | Major / labor-heavy | Monitor with a recorder first; do not gamble big labor on a hunch |
| Weak / clean tests | Low | Any | Monitor and document; replace only if it recurs with evidence |
| Pattern known, fault not yet caught | Any | Any | Monitor with instrumentation through one trigger window |
When to replace on suspicion (and do it honestly)
Replacing without proof is sometimes the right call, but only when you say so plainly. Use a suspicion-replace when the part is low-cost and quick, the consequence of another failure is high, and one component is clearly the most likely even if it passed at rest. Tell the customer in these words: this part is the most probable cause, it is not proven, here is what it costs to try it, and here is what we do next if the fault comes back. A shared bet is fair. A silent bet you present as a diagnosis is not.
When to monitor (and make it real monitoring)
Monitoring is not "call me if it happens again" with nothing left behind. Real monitoring means:
- A logging meter, min/max recorder, or memory clamp left on the suspect line so the next fault is captured.
- The customer briefed to record the exact time and conditions of the next failure.
- A written note of what passed and what you are watching, so the return visit starts with data, not a blank page.
If you cannot leave instrumentation and there is no pattern to wait out, at least set the customer's expectation honestly: the fault did not show, swapping good parts blind would waste their money, and the smart move is to catch it in the act.
Decide and recap
- One part moved the reading - replace it, then verify the fault is gone.
- Clean tests, high consequence, cheap part - replace the likeliest, instrument, watch.
- Clean tests, high consequence, expensive part - monitor with a recorder before betting big labor.
- Clean tests, low consequence - monitor and document.
- Known pattern - instrument and wait out the trigger.
The judgment to bank: never replace a major part on a hunch, and never call "monitor" without leaving something behind to catch the fault. Match the bet to the evidence and the stakes, and tell the customer which bet you are making.
References
- Trade-standard practice for intermittent-fault isolation and conditional monitoring
- Manufacturer documentation on component tolerances and acceptable readings
- OSHA lockout/tagout and energized-work guidance for any hazard branch (29 CFR 1910)
- See related: The Fault That Only Happens When You're Gone; The Ghost Fault: Document and Monitor