How to Log an Intermittent Fault Over Days
Why this matters
You cannot camp on a customer's site for a week waiting for a fault to show, and the fault knows it, it waits until you leave. A disciplined log, run by the customer or an instrument over days, captures the occurrences you will never personally witness. Done well, it turns "it just does that sometimes" into a page of dated events with conditions attached, which is enough to diagnose. Done badly, you get "yeah it happened a few times" and you are no further ahead.
Step 1: Decide manual log, instrument, or both
Match the tool to the fault. A visible, obvious event (a noise, a trip, a leak the customer sees) suits a manual log. A fast or invisible event (a voltage sag, a brief pressure spike, an overnight temperature swing) needs an instrument that records on its own. Often you run both: the customer notes what they experience, the logger captures what they cannot.
Step 2: Design a log sheet that captures what matters
A blank "write down when it happens" note is close to useless. Give them fields:
- Date and time (as exact as they can manage)
- What was running or in use at the moment
- Conditions (weather, who was home, what they had just turned on)
- What the fault did (severity, how long it lasted)
- What they did about it (reset it, waited, called)
Step 3: Brief the customer so it actually gets filled in
Make it dead simple and show them once. Tape the sheet where the fault happens, with a pen. Tell them plainly that a fault they do not log is a fault you cannot fix, and that a few complete entries beat a page of "it did it again." Set the expectation that catching the conditions is the whole point.
Step 4: Set the run length and a check-in
Decide how long the log runs before you return, based on how often the fault appears. A daily fault needs only a few days; a weekly one needs a few weeks. Set a check-in partway through, so a log that is going wrong (blank, or missing the key fields) gets corrected before you have wasted the whole window.
Step 5: Read the completed log for a pattern
When you get it back, treat it like evidence. Line the entries up and look for what every occurrence shares: a time, a load, a weather condition, a recent action. Feed it into a timeline. The shared condition across entries is your trigger, and if two conditions always appear together you may need one more round to tell which one matters.
Step 6: Know when you have enough to act
Stop logging when the entries point clearly at one trigger, or when they rule out the ones you suspected. A log that has captured the same conditions on every occurrence has done its job; more entries will not sharpen it. If weeks of logging show no pattern at all, that itself is a finding, the fault may be truly random hardware, and it is time to load-test the suspects rather than keep watching.
References
- Trade-standard practice for customer-run fault logging and observation
- See related: How to Set Up a Monitor to Catch a Fault You Can't Stay For
- See related: How to Build a Timeline of When a Fault Appears