You Only Get One Visit to Diagnose an Intermittent (Decision Tree)
Why this matters
An intermittent fault is hard enough when you can come back and watch it. Sometimes you cannot: the site is remote, the customer will not pay for a stakeout, access is a one-time window, or the unit is booked solid and cannot be tied up. You get one visit, the fault is not happening while you are there, and you still have to leave with something defensible. Winging it here means either replacing a part on a guess or telling the customer "call me when it does it again," which is not a diagnosis. This tree is how you spend a single visit to get the most diagnostic value the constraint allows.
Start here: budget the visit like the scarce resource it is
You have limited time and no second chance, so decide up front what each block of the visit is worth. The order that pays off most: interview, then capture the current scene, then attempt reproduction, then set up to catch it after you leave. Do not burn the whole visit on reproduction attempts that may never trigger and leave no time for the interview and the leave-behind, which produce value whether or not the fault shows.
If the intermittent involves a real hazard (arcing, a gas or CO concern, a protective device that trips), treat the hazard as present even though it is absent right now. An intermittent that is dangerous when it happens gets the safety response, not a "monitor and see," and that can mean shutting the equipment down until it is proven safe.
The interview does the observing you cannot
When you cannot watch the fault, the person who lives with it is your instrument. Interview for the condition, not the complaint.
- When does it happen: time of day, weather, after a specific use, when another load runs, hot start vs cold start.
- How often, and is it trending: getting more frequent points to a degrading part; steady and rare points to a specific trigger.
- What makes it stop: cycling power, waiting, tapping something. That tell narrows the fault type.
- What changed just before it started: any repair, install, or new appliance nearby.
Pin answers to specifics. "It acts up in the afternoon when the AC is on" is a lead; "sometimes" is not.
Capture the scene while you have it
Even with no fault present, the healthy state is worth recording. Take your key readings now, in the normal condition, so a later comparison has a baseline. Photograph connections, note anything marginal, and record the values you would want to compare against if you could see the fault. A baseline captured today is what makes tomorrow's fault data mean something.
Try to force it, within the time you budgeted
Spend your reproduction block deliberately, driving toward the condition the interview pointed at: load it, heat or cool it, run the full cycle, wiggle-test suspect connections, run the neighbor load the customer named. If it triggers, you have a live fault and a normal diagnosis. If your budgeted time runs out, stop and move to the leave-behind rather than chasing it indefinitely.
Set it up to catch itself after you leave
This is the highest-value use of a single visit, because it keeps working when you are gone.
- Leave a data logger or recording device on the suspect signal if you have one, so the fault records itself with a timestamp.
- Instrument the customer: give them a short, specific list of what to note the instant it happens (time, what was running, what the display showed) and how to capture it, so their next occurrence becomes usable evidence.
- Set a marker where practical, so the next event leaves a trace you can read on a return or that the customer can report.
The defensible call when you cannot confirm
If the visit ends with no reproduction, you still owe the customer a decision, honestly framed.
- Instrumented monitor is the right call when the fault is rare, not hazardous, and a logger or the customer's capture can plausibly catch it. You are buying a confirmed diagnosis before spending on parts.
- Empirical replacement of the single highest-probability suspect is defensible when the interview and baseline point hard at one part, the part is a reasonable cost against another trip, and you tell the customer plainly that it is the strongest lead, not a confirmed fix, and what the next step is if the fault returns. Do not present a probability as a certainty, and do not fire the parts cannon at several suspects at once.
The judgment to bank: one visit does not mean one guess. Interview, baseline, and a leave-behind turn a single shot into a real diagnosis or an honest, evidenced plan.
Recap
- Budget the visit: interview and leave-behind first, they pay off whether or not the fault shows.
- Treat an intermittent hazard as present even when it is absent.
- Use the customer as your observer and capture a healthy baseline while you can.
- Set the fault up to record itself, and if you still cannot confirm, choose instrumented monitor or a single honest empirical swap, never a multi-part guess.
References
- Trade-standard practice on intermittent-fault capture and baseline documentation
- Manufacturer documentation on data logging and monitoring accessories where available
- See related: The Fault Won't Happen While You're There; Monitor vs Replace on an Intermittent