Intermittent Fault You Cannot Reproduce On Site: Leave Vs Monitor Decision Tree
Why this matters
An intermittent fault that will not show itself while you are on site is the diagnostic situation most likely to produce a wasted callback, an unfair "you didn't fix it" complaint, or a guessed-at parts replacement that does not solve anything, because you cannot fix what you cannot observe. The pressure is to do something so the trip is not wasted, and that pressure drives techs to replace a plausible part on speculation, which often fails to fix the problem and leaves the customer paying for a guess. The deciding factors are whether the fault is safety-related, whether you can capture data or set up monitoring, and whether replacing the most-probable component is justified or just hopeful. Handling this honestly, by setting expectations and instrumenting the problem rather than guessing, protects the customer from paying for shotgun repairs and protects you from the callback when the guess does not hold.
The situation
The customer reports a problem that is real but is not happening while you are there. Everything tests normal, the fault does not present, and you cannot duplicate the conditions that trigger it. You have to decide whether to leave with a plan and clear expectations, set up monitoring or data capture to catch the fault, or replace the most-likely component on a probability basis. The honest difficulty is that without observing the fault, any part replacement is a bet. The temptation to replace something so the visit produces a tangible action is exactly what leads to paying for a part that does not fix the problem.
What is actually at stake
Customer trust is the central stake: an intermittent that comes back after a guessed repair reads as incompetence or as upselling a part that was not needed. Second is the customer's money: replacing components on speculation can run up a bill with no resolution. Third is safety: some intermittent faults (combustion, electrical, anything that can fail dangerous) cannot be left to recur and must be addressed conservatively even without reproduction. Fourth is your own callback exposure and reputation, both of which suffer when you commit to a fix you could not verify. Fifth is the diagnostic record, because a captured fault on a return visit is worth more than a dozen guesses.
Decision factors
- Is the fault safety-related? A combustion, electrical, or other potentially-dangerous intermittent cannot be left to recur. Conservative action (de-rate, disable, or replace a suspect safety component) may be warranted even without reproduction.
- Can you capture data? Monitoring equipment, logging, or recording controls can catch the fault when it next occurs, turning a guess into an observation.
- Is one component the clear high-probability cause? Sometimes the symptom pattern points strongly to one part; a justified probability-based replacement differs from a shotgun.
- Can the customer help capture conditions? When it happens, what is running, time of day, temperature, what they were doing. The customer's log is a real diagnostic tool.
- What is the frequency and trend? A fault that is rare versus worsening changes the urgency.
The decision
- Leave with a plan and expectations. No safety angle, no clear single cause, and you can instrument it. Tell the customer honestly that the fault did not present, set up monitoring or ask them to log conditions, and schedule a return when the fault is caught. Do not replace parts on a hope.
- Monitor / capture data. The fault is reproducible enough to instrument. Install logging or monitoring, or coach the customer on what to record, so the next occurrence is observed rather than guessed.
- Replace the high-probability component. The symptom pattern points strongly to one part, the replacement is justified by the evidence (not just availability), and you disclose to the customer that it is a probability-based fix. Set the expectation that it may not be the final answer.
- Act conservatively for safety / escalate. The intermittent involves a potential safety failure. Take conservative protective action even without reproduction, and escalate if the safe path is unclear. A safety intermittent is not something to leave to recur.
Why the shotgun repair fails
The pressure to make the trip productive is what drives a tech to replace a plausible-looking part on speculation, and that shotgun approach fails the customer on every axis. It often does not fix the problem, because without observing the fault you are guessing at the cause, so the customer pays for a part and the intermittent comes back. It runs up the bill one hopeful component at a time when the fault recurs, which the customer experiences as either incompetence or as being upsold parts they did not need. And it poisons the diagnostic record, because now you cannot tell whether the replaced part mattered. The honest alternative is to tell the customer plainly that the fault did not present, that replacing parts blind would be a guess they should not pay for, and that the reliable path is to catch the fault in the act through monitoring or their own log. Customers respect a tech who refuses to charge for a guess far more than one who replaces a part and hopes, and a fault captured on the next occurrence is worth more than any amount of speculation.
What to document
Record the reported symptom and the conditions the customer says trigger it, everything you tested and the normal readings, the fact that the fault did not reproduce, the plan (monitoring, customer log, scheduled return, or probability-based replacement and the reasoning), and what you told the customer about expectations. If you replaced a part on probability, document that it was disclosed as such. This record is essential because the next tech, or you on the return, needs the history, and it protects you from an unfair "didn't fix it" claim on a fault you could not observe.
References
- ACCA and PHCC standards of practice on honest diagnosis and not representing a speculative repair as a confirmed fix.
- Manufacturer service literature and fault-code documentation for the specific equipment, which guides probability-based diagnosis of intermittent faults.
- NFPA 54 / National Fuel Gas Code and NEC / NFPA 70 where the intermittent involves combustion or electrical safety, which govern conservative protective action.
- Your company's diagnostic-documentation and callback policy.