One Failed Unit, No Second Data Point Yet, a Technical Read on Timing
Why this matters
You are standing in front of exactly one failed unit. There is no second instance yet, no fleet to search, no pattern to compare against, because you have not gone looking for one. The question is narrower than "is this a batch defect": given only this one failure and its timing, is there enough technical reason to suspect something other than ordinary wear, before you invest any time chasing a second data point. Answering that from a single unit is a teachable skill, and it is the gate that decides whether a multi-unit investigation is even worth starting.
The concept you need: the bathtub curve
Almost every component's failure rate over its life follows a recognizable shape, commonly pictured as a bathtub in cross-section: high failure risk at the very start, a long low-and-flat stretch in the middle, then rising risk again at the end.
- Infant mortality (the early wall of the tub). A cluster of failures very early in life, driven by a manufacturing defect, a damaged part, or an installation error that did not show up until the component was put under load. Failure risk here is elevated but drops quickly once the weak units fail and the survivors keep running.
- Random failure period (the flat bottom). The long stretch where failure risk is low and roughly constant, driven by random stress events rather than any systemic weakness. Most of a component's service life sits here.
- Wear-out period (the rising wall at the end). Failure risk climbs again as the component approaches and exceeds its expected service life, driven by cumulative fatigue, corrosion, insulation breakdown, or mechanical wear finally catching up.
A single failure's position on this curve is your first real clue, available with just one unit in hand.
Step 1: place the failure on the curve using actual service life, not calendar age
Get the real in-service time, not the manufacture date and not the install date alone. A unit installed and run hard immediately has a very different service-life clock than one left mostly idle for years before this failure.
- If the failure lands solidly in the wear-out region (at or past the expected service-life range for this component category, adjusted for how hard it ran): ordinary wear is the leading explanation. You do not need a second data point to close this out.
- If the failure lands in the flat, random-failure middle of the curve, well short of expected wear-out but past any reasonable infant-mortality window: this is a random stress failure, not systemic. Investigate the specific stress event (a surge, an overload, an environmental spike) rather than wear or a batch theory.
- If the failure lands very early, inside the infant-mortality window for this component type: this deserves the rest of this checklist, because early failure is the signature a manufacturing defect or an installation error leaves on a single unit.
Step 2: adjust the expected-life curve for environmental severity before you judge "early"
The same component in two different environments does not share the same expected-life curve, and skipping this adjustment is the most common way a tech misjudges a premature failure.
- Heat, vibration, contamination, moisture, and duty cycle all compress the curve. A component rated for a typical service life in a benign environment can reasonably wear out much sooner in a harsh one, and that shortened life is still ordinary wear, not evidence of a defect.
- Ask what this unit's actual operating conditions were, not the nameplate-rated conditions. A unit run in continuous heavy duty cycle, high ambient heat, or a corrosive atmosphere has a legitimately shorter wear-out expectation than the same part in a mild, intermittent-duty application.
- Only call a failure "early" after this adjustment. A failure that looks premature against a generic service-life number can turn out to be squarely on schedule once you account for how hard the environment worked it.
Step 3: distinguish infant-mortality failure signatures from wear-out failure signatures
Even without a second unit, the failure mode itself often tells you which side of the curve you are on, because the two failure types tend to look different.
- Infant-mortality signatures: a failure that traces to a defect present since manufacture (a flaw, contamination, or damage at the failure point that predates any real operating stress), or to an installation step done wrong from day one, rather than gradual degradation.
- Wear-out signatures: a failure that shows progressive degradation consistent with the time and stress the unit has actually seen: cumulative fatigue marks, gradual insulation breakdown, mineral or corrosion buildup matching its service duration, rather than a flaw that looks like it was always there.
A young unit with a wear-out-looking signature, or an old unit with an infant-mortality-looking signature, is a mismatch worth flagging even from a single instance.
What this checklist tells you, and what it does not
| Finding on one unit | What it means | What to do next |
|---|---|---|
| Failure lands in the wear-out region for this environment | Ordinary wear | Repair and close out, no further investigation needed |
| Failure lands in the flat middle, environment-adjusted | Random stress event | Investigate the specific stress, not age or a batch theory |
| Failure lands early, environment-adjusted, with a wear-out-looking signature | Ambiguous; recheck the environmental adjustment | Confirm actual duty cycle and conditions before concluding anything |
| Failure lands early, environment-adjusted, with an infant-mortality signature | Premature failure with a technical basis to suspect a manufacturing or install defect | Move to the multi-instance investigation |
This checklist only answers whether one unit gives you enough technical reason to suspect something other than wear. It does not tell you whether this is a batch defect, a design flaw across the product line, or an isolated fluke. That determination requires a second confirmed instance sharing the same manufacture window and failure mode, and a separate investigation once you have it. If this gate says "premature," that is the trigger to start the fleet-wide search, not the place to conclude a batch or design verdict from one unit alone.
Recap
- Get actual in-service time and duty cycle, not just calendar age.
- Place the failure on the bathtub curve: infant mortality, random middle, or wear-out.
- Adjust the expected-life range for this unit's real environmental severity before judging "early."
- Match the failure signature (flaw-like vs degradation-like) to infant-mortality or wear-out.
- If the gate says premature with a technical basis, hand off to the separate multi-instance batch-versus-design investigation; do not answer that question from one unit alone.
References
- Manufacturer documentation on expected service life and rated operating conditions
- Reliability-engineering literature on the bathtub curve and infant-mortality failure analysis
- Trade-standard practice for premature-failure investigation
- See related: One Bad Batch vs a Design Flaw: Decision Tree; Reading a Manufacture or Install Date Code for a Cohort-Wide Pattern