Fails Only on the Hottest or Coldest Day of the Year: Decision Tree
Why this matters
A system that only fails on the single hottest day or the single coldest night of the year is one of the most frustrating calls to diagnose, because by the time you arrive the extreme condition has usually passed and the equipment is running fine again. Chasing this with tests run under normal conditions almost always comes back clean, and a tech who declares "no fault found" on a genuinely marginal system sets the customer up for the exact same failure next time the temperature spikes again. This tree is about finding a fault that is real but only reveals itself at the edge of the equipment's operating range.
Start here: is this a capacity problem or a component problem
Before you test anything, separate two very different causes that produce the identical symptom.
- The equipment is undersized or degraded relative to the actual peak demand, and simply cannot keep up once the load crosses a threshold. Everything about the system can be functioning correctly and it still falls behind at the extreme.
- A specific component is marginal and fails under peak electrical, thermal, or mechanical stress, while the rest of the system is fine. This is a true fault, just one that only manifests under conditions you cannot recreate on a mild day.
Ask the customer precisely what "fails" meant: did it run continuously but not keep up (capacity), or did it shut down, trip, or stop entirely (component fault)? That single distinction often points you toward one branch or the other immediately.
If it looks like a capacity problem
- Check the equipment's rated capacity against the actual peak demand of the space or application. A system sized correctly for typical conditions can still be undersized for the statistical extreme; this is a sizing and expectation conversation, not a repair.
- Check for degradation that reduces effective capacity below its original rating - buildup, wear, reduced efficiency, a component operating below its original output even though it has not failed outright. A system that once handled the extreme fine and now cannot may have quietly lost capacity over time rather than having always been marginal.
- If the system was always marginal for the true extreme, the honest answer is that this is a capacity limitation, not a defect, and the fix is an upsize, a supplement, or an expectation reset with the customer, not a parts replacement.
- If capacity has degraded from a prior baseline, diagnose what caused the loss (see the related articles on reading wear patterns and on connected-system data logs if the equipment logs performance over time) and address that cause, which may restore enough margin to handle the extreme again.
If it looks like a component-level fault
- Identify which specific component is stressed hardest at the temperature extreme, not just at normal running conditions. Extreme heat often stresses electrical components (increased current draw, elevated internal temperatures pushing a marginal part over its limit); extreme cold often stresses mechanical components (increased fluid viscosity, contraction-related clearance issues, slower or harder starts).
- Check the component's rating against the actual extreme, not the average operating condition. A part rated with little margin above typical peak load is exactly the kind of thing that survives all season and fails on the one day that exceeds its rating.
- If you cannot test at the actual extreme condition, look for secondary evidence instead: discoloration or heat-stress marks on an electrical component, unusual wear concentrated at one point, a connected system's data log showing a reading trending toward its rated limit during the reported event (see the related article on reading a connected system's data log).
- Consider recreating a proxy condition if it is safe and practical - artificially loading the circuit, restricting airflow briefly under controlled and monitored conditions, or otherwise pushing the system toward its limit in a controlled way rather than waiting for the next true extreme. Only do this within the equipment's safe test parameters; forcing a component past a rating you have not confirmed is safe to exceed is not a diagnostic step, it is a way to cause damage.
- If you find a component that is marginal but has not yet definitively failed, you are making a probabilistic call: replace preemptively based on strong circumstantial evidence, or wait for a harder failure to confirm. Be honest with the customer about which one this is; "I'm confident but not certain" is a legitimate thing to say when the evidence is strong but the extreme condition could not be directly reproduced.
When you find nothing and the customer's account is credible
A system testing clean under normal conditions does not mean nothing is wrong; it may mean the fault genuinely only exists at the extreme. Rather than closing the call as no-fault-found, consider: setting up temporary monitoring that will capture the next extreme event, scheduling a follow-up specifically timed near the next forecast extreme, or, if the equipment has connected logging, reviewing historical data from prior extreme days for a pattern you can act on now instead of waiting for the next one.
Recap
- Determine whether the failure is a capacity limit or a component fault.
- For capacity, compare rated capacity (and any degradation) against actual peak demand.
- For a component fault, identify what is stressed hardest at the extreme and check its rating against that real condition, not the average.
- When you cannot reproduce the extreme directly, use secondary evidence, monitoring, or a timed follow-up rather than closing the call as no-fault-found.
References
- Manufacturer documentation on rated capacity and component operating limits at temperature extremes
- Trade-standard practice for load-related and capacity-related fault diagnosis
- See related: The Seasonal First-Use Fault, Different From a Transition Fault; Reading a Connected System's Data Log: The Discipline