Symptom Gone After Customer Tinkered: Investigate vs Leave Decision Tree
Why this matters
The "it was doing it yesterday but it's fine now" arrival is one of the highest callback-risk conditions in field service. The customer poked at something, the symptom vanished, and the technician faces a choice: spend billable time hunting a fault that is not currently present, or document the visit and leave. Both options have failure modes. Leaving without investigation invites a second dispatch within days, often at no charge under warranty or goodwill. Over-investigating a system that is now operating to specification burns time the customer will resist paying for and can itself induce new faults. ISO 14224 frames this as the distinction between latent and active failures, and the same logic applies whether the asset is a furnace, a pool pump, a refrigerator, or a roof leak. The decision tree below structures the call so the technician leaves with either a confirmed repair, a confirmed monitoring plan with measurable triggers, or a documented no-fault-found that will not embarrass the company on the next call.
Step 1: Reconstruct what the customer actually did
Before touching the equipment, get a precise narrative. Open-ended questions ("walk me through exactly what you did, in order") produce better recall than yes/no prompts. Anchor the timeline: when did the symptom start, when did the customer intervene, when did it stop, how long has it been fine. Ask what specifically was touched: a breaker, a switch, a valve, a setting, a filter, a reset button. Ask whether anything was added: a chemical, a cleaning product, a new device on the circuit, a replacement part. Per ACCA Quality Installation guidance, undocumented customer interventions are the leading cause of recurrent service tickets that read as "intermittent" but are actually deterministic responses to an unrecorded change. If the customer cannot reconstruct the sequence, treat the call as a latent fault with unknown trigger and proceed to Step 3.
Step 2: Classify the intervention
Customer interventions fall into four categories, each requiring a different response.
Resets and power cycles. Breakers flipped, GFCI reset, equipment unplugged and replugged, thermostat batteries pulled. A symptom that disappears on power cycle and stays gone for hours points at a transient lockout (overheat, overcurrent, communication timeout) or a soft-fault microcontroller state. The fault will return; the question is when.
Cleaning and clearance. Filter changed, drain cleared, vent unblocked, coil rinsed, strainer emptied. If the symptom was a flow or airflow restriction and the cleaning removed it, the root cause is the restriction source, not the asset. Document what was clogged and quantify the next-cleaning interval.
Setting changes. Thermostat schedule altered, pool timer reset, water heater setpoint changed, mode switched. A symptom that resolves on a setting change is not a fault; it is a configuration mismatch. Confirm the setting is now correct and educate.
Chemical or consumable additions. Drain cleaner, water-treatment chemicals, refrigerant top-off, sealant. These can mask a leak or restriction temporarily. Treat as high-risk for recurrence.
Step 3: Run a non-intrusive baseline
Regardless of category, capture the current operating state before deciding to dig deeper. Measure what is measurable without disassembly: voltages at the supply and at the load, current draw under normal cycle, temperatures across the relevant interfaces, pressures, flow indicators, error-code history if the equipment logs one. Photograph nameplate and any visible wear. This baseline serves two purposes. First, it establishes whether the system is currently within manufacturer specification or merely "running" while drifting. Second, it gives the next technician a reference point if the customer calls back. NFPA 70B Chapter 9 codifies this practice for electrical equipment as condition-based maintenance baseline capture, and the same discipline applies across trades.
Step 4: Decide using the four-quadrant matrix
Plot the intervention against the baseline.
Quadrant A: Intervention was a reset, baseline shows readings within spec, no error history. The fault is intermittent and not currently observable. Document the visit as monitoring, set explicit return triggers (specific error codes, specific symptoms, specific noises), and leave. Do not replace parts on suspicion alone.
Quadrant B: Intervention was a reset, baseline shows a reading at the edge of spec or error history present. The fault is latent and will recur. Investigate the suspect subsystem now. The customer's reset bought you a working system to test against, which is more useful than a dead one.
Quadrant C: Intervention was cleaning or configuration, baseline is in spec. The symptom was caused by the dirty filter or wrong setting. Confirm root cause was the restriction or the setpoint, document, and leave with a maintenance interval recommendation.
Quadrant D: Intervention was chemical or consumable, baseline is in spec. The symptom may be masked. Investigate the system that the additive treated. A pool that the customer shocked, a drain the customer poured cleaner into, an HVAC system the customer topped off with refrigerant. None of these are repairs; they are stalls.
Step 5: Set explicit return-trigger language
If the decision is to leave without further work, the documentation must give the customer specific, observable triggers for a re-dispatch. Vague guidance ("call us if it acts up again") generates ambiguous callbacks. Specific triggers ("call us if the error code 7 returns, if you hear the same rattle for more than 30 seconds, if the temperature drops below 68 in heating mode") produce either a clean callback with diagnostic data or no callback at all. PHCC and NARI service-standard guidance both emphasize that customer-observable triggers are a service-quality control, not a cop-out. Write them on the invoice.
Step 6: When to escalate to deeper diagnosis
Escalate past the baseline regardless of quadrant when any of the following apply: the equipment is under warranty and a no-fault-found will not be reimbursable; the symptom involves a safety-rated control (limit switch, pressure relief, GFCI, smoke alarm) where intermittent operation is itself a failure mode; the asset is high-criticality (commercial refrigeration, server-room cooling, life-safety) where the cost of recurrence exceeds the cost of investigation; the customer's intervention defeated a protective function and the function has not yet been verified reintroduced.
References
- ISO 14224:2016, Petroleum, petrochemical and natural gas industries, Collection and exchange of reliability and maintenance data for equipment, Annex B on failure mode classification.
- NFPA 70B-2023, Recommended Practice for Electrical Equipment Maintenance, Chapter 9, Condition-Based Maintenance.
- ACCA Standard 5, HVAC Quality Installation Specification, Section on documentation and customer-induced conditions.
- OSHA 29 CFR 1910.147, Control of Hazardous Energy, applicable when intervention involved disabling a protective control.
- PHCC National Standard Plumbing Code, service documentation guidance.