Documenting a Cohort Pattern Well Enough to Act on It
Why this matters
"We keep seeing these fail" is a hallway comment. The exact same observation, written down with the right fields, is a report a supplier acts on, a warning you can send to every affected customer, and a defense if a failure ever ends up in a dispute. The gap between the two is not the observation, it is the documentation. Most shops that notice a real cohort pattern (a group of units sharing an age, a batch, an installer, or a site condition) never capture it well enough to do anything with it, and the pattern dies as a shared memory that fades in a few months. This is the record-keeping discipline that turns a hunch into evidence.
What "well enough to act on" actually requires
A pattern report is actionable when someone who was not there, reading it cold, can answer three questions from the document alone: what exactly failed, which units share the risk, and what evidence proves this is a real pattern rather than a coincidence. If your notes cannot answer all three, they are not documentation yet, they are a memory aid for you personally.
The fields that make a pattern report actionable
Capture these for every instance, not just the first one you notice:
- Exact failure mode. Not "it broke," but the specific component and how it failed (cracked, seized, shorted, corroded through). Vague failure descriptions are the single biggest reason a supplier or manufacturer dismisses a pattern report; "it stopped working" could be anything.
- Identifying information for the unit. Type, manufacture date or date code if visible, install date, and any batch or lot marking. This is what lets you define the cohort, the actual set of units that share the risk, rather than a vague "units like this one."
- Age or run time at failure. A component failing at a strikingly consistent, early point across multiple units is much stronger evidence than the same failure scattered across a wide range of ages.
- Site and install context. Water quality, power quality, climate or environmental exposure, and who installed it, if known. This is what lets you tell a batch-defect pattern from an environmental or installation-practice pattern, which need completely different fixes.
- Count and date range of instances. How many, over what window. A pattern report with a real count and a real date range reads as evidence; one that says "several" and "recently" reads as an impression.
Build the cohort definition, not just the failure list
The single most useful thing a good pattern document produces is a clear answer to "who else is affected." That means defining the cohort precisely: not "everyone with this type of unit," but "units of this type, installed in this date range, from this batch or by this installer, showing this specific failure mode." A precise cohort definition is what lets you proactively contact the right five customers instead of alarming fifty who were never at risk, or missing the ones who actually are.
Three or more before you call it a pattern
Two similar failures can be coincidence. Three to five, sharing a specific factor (a manufacture date window, an installer, a site condition), is where a pattern claim starts to carry real weight, though the threshold moves with how unusual the failure mode is: a genuinely rare failure mode showing up twice, in a way you have never seen before, deserves documentation and a cautious flag even below that count. Document every instance from the first one regardless of the eventual count, since you will not know you are looking at a pattern until you have enough of them, and by then the details of the first one may be a fading memory if you did not write them down at the time.
Format the report so it survives being forwarded
A pattern report someone else can act on should read cleanly even out of context, because it will likely get forwarded to a supplier, a manufacturer, or another location in your own company:
- Lead with a one-line summary: the failure mode, the count, and the date range.
- Follow with the cohort definition: exactly which units share this risk.
- List the instances in a simple table: date, unit identifier, failure mode, age or run time at failure.
- Close with what you are asking for: an investigation, a known-issue check, a proactive parts hold, or simply an acknowledgment that the pattern is on record.
A report in this shape gets read and acted on. A narrative paragraph describing "the recent run of failures we've noticed" gets skimmed and set aside.
Keep the record even if nothing comes of it yet
Not every documented pattern resolves into a confirmed batch defect or a manufacturer bulletin. Keep the record anyway. The fourth instance six months from now, showing up after you had stopped actively tracking, is far easier to connect to the first three if you wrote them down properly the first time.
References
- Trade-standard practice for recurring-failure investigation and reporting
- Manufacturer documentation on batch and serial-date tracking conventions
- See related: Reading a Fleet-Wide Failure Pattern
- See related: One Bad Batch vs a Design Flaw Decision Tree