How to Log an Intermittent Fault Over Days

Why this matters

You cannot camp on a customer's site for a week waiting for a fault to show, and the fault knows it, it waits until you leave. A disciplined log, run by the customer or an instrument over days, captures the occurrences you will never personally witness. Done well, it turns "it just does that sometimes" into a page of dated events with conditions attached, which is enough to diagnose. Done badly, you get "yeah it happened a few times" and you are no further ahead.

Step 1: Decide manual log, instrument, or both

Match the tool to the fault. A visible, obvious event (a noise, a trip, a leak the customer sees) suits a manual log. A fast or invisible event (a voltage sag, a brief pressure spike, an overnight temperature swing) needs an instrument that records on its own. Often you run both: the customer notes what they experience, the logger captures what they cannot.

Step 2: Design a log sheet that captures what matters

A blank "write down when it happens" note is close to useless. Give them fields:

  • Date and time (as exact as they can manage)
  • What was running or in use at the moment
  • Conditions (weather, who was home, what they had just turned on)
  • What the fault did (severity, how long it lasted)
  • What they did about it (reset it, waited, called)

Step 3: Brief the customer so it actually gets filled in

Make it dead simple and show them once. Tape the sheet where the fault happens, with a pen. Tell them plainly that a fault they do not log is a fault you cannot fix, and that a few complete entries beat a page of "it did it again." Set the expectation that catching the conditions is the whole point.

Step 4: Set the run length and a check-in

Decide how long the log runs before you return, based on how often the fault appears. A daily fault needs only a few days; a weekly one needs a few weeks. Set a check-in partway through, so a log that is going wrong (blank, or missing the key fields) gets corrected before you have wasted the whole window.

Step 5: Read the completed log for a pattern

When you get it back, treat it like evidence. Line the entries up and look for what every occurrence shares: a time, a load, a weather condition, a recent action. Feed it into a timeline. The shared condition across entries is your trigger, and if two conditions always appear together you may need one more round to tell which one matters.

Step 6: Know when you have enough to act

Stop logging when the entries point clearly at one trigger, or when they rule out the ones you suspected. A log that has captured the same conditions on every occurrence has done its job; more entries will not sharpen it. If weeks of logging show no pattern at all, that itself is a finding, the fault may be truly random hardware, and it is time to load-test the suspects rather than keep watching.

References

  • Trade-standard practice for customer-run fault logging and observation
  • See related: How to Set Up a Monitor to Catch a Fault You Can't Stay For
  • See related: How to Build a Timeline of When a Fault Appears