Intermittent Fault You Cannot Reproduce: Monitor vs Replace vs Return Decision Tree
Why this matters
An intermittent fault that will not show itself while you are standing in front of the equipment is one of the most expensive situations in field service, because every wrong move costs a return trip. Replace a part on a hunch and the symptom recurs, you have burned a part and your credibility while solving nothing. Walk away saying you found nothing, and you look incompetent when it acts up an hour later. The customer is paying for certainty you genuinely do not have yet, so the real skill is choosing honestly among monitoring, evidence-based replacement, and a scheduled return without pretending to a diagnosis you cannot confirm. Handled as a deliberate decision with clear documentation, an intermittent fault becomes a managed process instead of a string of frustrating callbacks.
The situation
The customer describes a problem that happens sometimes: a unit that occasionally trips, a noise that comes and goes, a fixture that fails once a day. On arrival everything works. You cannot make it fail on command, and the customer cannot make it fail for you either. You have to decide what to do with a system that is currently behaving normally but is reported to misbehave when you are not watching. The customer is standing there expecting a verdict, and the temptation to manufacture one, by pointing at the most likely-looking part and replacing it, is strong precisely because saying "I could not reproduce it" feels like failure. Resisting that temptation and replacing it with a structured plan is the entire skill here.
What is at stake
The core risk is the return-trip spiral, where you replace one suspect part, leave, get called back, replace another, and never converge because you are guessing. Each loop costs labor, parts, and trust. There is also a liability dimension: if the intermittent fault is a safety item (an electrical fault, a gas-related anomaly, a brake-like failure in equipment), leaving it unresolved while telling the customer it is fine transfers risk onto you. Conversely, replacing expensive parts the customer pays for without evidence invites a billing dispute when the symptom persists. Honesty about diagnostic uncertainty, documented clearly, is your protection on both fronts.
Decision factors
- Quality of the evidence trail. A fault that leaves a record (an error code in memory, a tripped device, scorch marks, a logged event) is far more diagnosable than one with no trace.
- Reproducibility under stress. Can you provoke it by loading the system, cycling it, heating or cooling it, or running it longer than your visit allows.
- Safety classification. An intermittent symptom that could indicate a fire, shock, gas, or structural hazard changes the calculus from convenience to risk and may force conservative action.
- Cost and certainty of the suspected part. Replacing a cheap, commonly-failing component on strong circumstantial evidence is reasonable; replacing a costly part on a guess is not.
- Customer tolerance and availability. Whether the customer can live with the symptom while you monitor, and whether they can capture data (photos, video, timestamps) between visits.
- Pattern correlation. Whether the customer can already tie occurrences to a trigger such as time of day, weather, a specific appliance running, or a particular load, since a correlation you can act on shrinks the search dramatically.
Options and when each wins
Monitor when the fault is not a safety item, leaves a usable record, and the customer can tolerate it briefly. Install or enable logging, set a follow-up window, and ask the customer to document each occurrence with time and conditions. This wins because it gathers the evidence that converts a guess into a diagnosis, and it costs only a small amount of equipment and a scheduled callback rather than a string of speculative parts. Replace on evidence when a specific component is a known high-probability cause for that exact symptom and the part is low-risk relative to a return trip; you are making a defensible bet, not a wild one, and you tell the customer it is a probable fix, not a guaranteed one, so that a recurrence does not feel like a broken promise. Schedule a return when you need conditions you cannot create on this visit (a specific time of day, a weather state, a load you cannot apply), and you set that expectation up front rather than leaving and hoping; returning under the failing condition is often the only way to catch a fault that hides during a midday visit. For any intermittent fault with a credible safety dimension, conservative action wins: de-energize, isolate, or take the equipment out of service and document why, rather than leaving a possible hazard live. The blended approach that resolves most intermittent faults cleanly is to leave monitoring in place, give the customer a simple way to capture each occurrence, and book the return for the condition under which it fails, so the next visit arrives armed with data instead of repeating the same dead-end inspection.
What to document
Record exactly what the customer reported, the conditions under which it occurs, and that you could not reproduce it on site. Note any captured evidence and any logging you left behind. If you replaced a part on probability, write that it was an evidence-based likely fix, not a confirmed repair, and that the customer was told the symptom may persist. If you chose to monitor or return, record the follow-up plan and the customer's agreement to it. For a safety-related decision to take equipment out of service, document the hazard and the customer's notification clearly.
References
- Occupational Safety and Health Administration, General Duty Clause, 29 U.S.C. 654(a)(1) (employer obligation regarding recognized hazards, relevant when an intermittent fault is a safety item)
- Magnuson-Moss Warranty Act, 15 U.S.C. 2301 et seq. (governs how replaced-part and workmanship representations must be honored)
- Air Conditioning Contractors of America (ACCA), diagnostic-procedure and service-documentation guidance
- Electronics Technicians Association International (ETA), field-diagnostic best-practice resources for intermittent faults